Research questionHow can LLMs generate reliable, adaptive tests that expose one another’s model-specific weaknesses?Static benchmarks become stale and may miss failures specific to individual models. Model-generated tests can also be inconsistent or lack trustworthy reference answers, making genuine weaknesses difficult to distinguish from flawed probes.