Get Started
Home
Topics
Search
Library
Research questionHow can we certify LLM-generated, sample-level hypotheses without circular verification or spurious correlations?LLMs can generate plausible but hallucinated hypotheses, while checking them with the same model can be circular. Held-out testing may also accept false hypotheses when apparent support arises from spurious correlations.
AI
Evaluation & Benchmarks
Machine Learning
Natural Language Processing
Statistical Machine Learning
Latest papersRecent research connected to this question, newest first.NxN E-valuation: Hypothesis Certification via a Conformal CRT NullThe source describes an e-value-based conformal conditional randomization approach that uses different samples from a large training set as null hypotheses for one another. Its stated applicability is limited to LLM generations that express hypotheses applying to each individual sample.research paper · Sep 3, 2026
Related questions
How can high-stakes LLM systems distinguish unsupported claims from novel ones and prioritize expert verification?How can LLMs produce reliable confidence estimates for deciding when to defer outputs to humans?How can watermarking provide trustworthy provenance for LLM text at scale despite transformations and accumulating false positives?How can human reviewers reliably detect LLM errors when verification reasoning is hard to retrieve?