Glenn Matlin
  • Home
  • About
  • Research
  • Work With Me
  • Reading Lists
  • Publications
  • Guides
  • Blog
  • CV

Role Steering of Language Models for Social Simulations

social simulation
activation steering
personas
An activation-steering screening workflow for role-conditioned agents in social simulations, with CastVectors role directions and a per-role coefficient screen.
Authors

Isaac Song

Mohammed Rehan Parwani

Glenn Matlin

Emile Anand

Akhil Theerthala

Arjun Chatterjee

Anthony Wen-Ming Zang

Maria Kostylew

Yonadav G. Shavit

Sebastien Krier

Mark Riedl

Published

August 8, 2026

Publication

Role Steering of Language Models for Social Simulations

An activation-steering screening workflow that checks role-conditioned agents before they are placed into a simulated population, with CastVectors role directions and a per-role coefficient screen.

Accepted

August 2026

Authors

Isaac Song†, Mohammed Rehan Parwani†, Glenn Matlin†, Emile Anand, Akhil Theerthala, Arjun Chatterjee, Anthony Wen-Ming Zang, Maria Kostylew, Yonadav G. Shavit, Sebastien Krier, Mark Riedl († equal contribution)

Venue

Social Simulation with LLMs: Fidelity in Applications Workshop at COLM 2026

All Publications

Abstract

Social simulations built from language-model agents need role-conditioned behavior that can be checked before agents are placed into a simulated population. We introduce an activation-steering screening workflow for role-conditioned agents: define a role profile, extract a role-specific direction, sweep four steering coefficients, evaluate role-profile alignment, and pass or flag each candidate configuration. On OLMo-3-7B-Instruct, we apply the workflow to a mixed 275-role inventory with 228 role-agnostic questions, GPT-4.1-mini prompted role references, and GPT-4.1-mini judges. We also introduce CastVectors, which receive higher judged role-profile alignment than an assistant-axis directional control from prior persona-vector work, with mean overall scores of 63.2 versus 41.1 across the tested grid. They also preserve high lexical diversity, while the control drops sharply at larger coefficients. The role-level screen is the main contribution: most roles improve as steering increases, but 38 roles decline across all six measured dimensions, showing why simulation builders should choose coefficients per role rather than deploy a uniform high-strength setting. We make our code and evaluation artifacts available at github.com/eilab-gt/casting-call-vectors.

At a Glance

  • A screening workflow for role-conditioned agents: role profile, role-specific direction, coefficient sweep, alignment evaluation, pass or flag
  • Applied to a 275-role inventory with 228 role-agnostic questions on OLMo-3-7B-Instruct
  • CastVectors score 63.2 versus 41.1 mean role-profile alignment against an assistant-axis control, while preserving lexical diversity
  • 38 roles decline across all six measured dimensions, showing coefficients should be chosen per role rather than uniformly

Cite This Paper

BibTeX
@inproceedings{song2026rolesteering,
  title     = {Role Steering of Language Models for Social Simulations},
  author    = {Song, Isaac and Parwani, Mohammed Rehan and Matlin, Glenn and Anand, Emile and Theerthala, Akhil and Chatterjee, Arjun and Zang, Anthony Wen-Ming and Kostylew, Maria and Shavit, Yonadav G. and Krier, Sebastien and Riedl, Mark},
  year      = {2026},
  booktitle = {Social Simulation with LLMs: Fidelity in Applications Workshop at the 3rd Conference on Language Models (COLM)}
}

Continue exploring

Return to the publication archive or step back to the broader research agenda.

Publications Research

© 2025-2026 Glenn Matlin