Get Started
Research questionHow can dangerous capabilities in LLMs be compared consistently across models and releases?Fragmented safety evaluations make it difficult to determine whether models differ in dangerous knowledge, resistance to unsafe requests, or harmful outputs. They also obscure whether newer releases are becoming safer.
AI
Alignment & Safety
Evaluation & Benchmarks
Latest papersRecent research connected to this question, newest first.FUSE: An Evaluating Framework for Dangerous Capabilities of LLMsThe evidence covers 12 commercial LLMs from four model families, with a chemical-biological evaluation and a cyber pilot. It separates knowledge, defense, and harm and tracks them by release date; reliability evidence comes from cross-judge consistency and inter-pipeline correlations, so conclusions remain bounded by the tested models, domains, and protocol.research paper · Sep 2, 2026
Related questions
How should LLM safety be evaluated when harmful prompts vary in implicitness and sophistication?How can LLM-agent systems prevent safety compromises from propagating across workflow boundaries?How can we evaluate LLM reasoning quality beyond final-answer accuracy across deployment contexts?When does improving LLM trader capability increase market risk through correlated decisions?
Home
Topics
Search
Library