Research questionHow can black-box LLM APIs detect hallucinations without trusted context while limiting false alarms?Without a reference document, an LLM may produce consistently confident errors or varied answers whose meaning is difficult to assess. Detection must therefore distinguish these failure patterns while avoiding excessive referrals for human review.