Research questionHow can we validate LLM policy simulations against observed actors, actions, and impacts?LLM policy simulations can produce plausible actor behavior without evidence that their predicted actions and downstream effects match observed records. Without that grounding, it is difficult to determine whether multi-agent interaction improves policy analysis or merely makes explanations more elaborate. Latest papersRecent research connected to this question, newest first.GPS-Bench: A Governance Policy Benchmark for Automating Policy AnalysisThe benchmark links policies to dated public evidence from legislative records, lobbying disclosures, regulatory documents, corporate filings, economic data, and related sources. Actors are reconstructed as provenance-bearing evidence objects, with human-annotated Gold cases and separately LLM-labeled Silver cases used only for supervision. The supplied comparisons use a common grounded policy state and include joint, independent, and communicating actor agents, graph-based methods, and weight-level fine-tuning, with predictions covering actor impacts and interaction mechanisms.research paper · Sep 3, 2026