Story

arxiv_cs_cl ยท Oct 6, 2026 ยท paper

Source brief

Latent space bias directions in LLMs capture confidence, not fairness

arxiv.orgOct 6, 2026
original source linked

In brief

Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models. However, prior work reports that steering vectors generalise poorly, with unintended effects on...

Feed lens
eval

Continue reading

Read the original at arxiv.org โ†’Open in live feedRead that dayโ€™s brief

Earlier in this thread 4 items