Get Started
Home
Topics
Search
Library
7 min read · Agents · Evaluation · Added Oct 8 · Paper published Oct 5, 2026

MiniCorp: The Last Mile of the AI Agent Firm

Source: research paper via Hugging Face Daily Papers
MiniCorp tackles the lack of counterfactual, long-horizon enterprise data for training agents to run companies, not finish tasks. A checkpointable e-commerce simulator pairs a persistent agent firm with an evolving market, letting the same week replay under different decisions — reproducing a 71% promo lift and 84% ad cannibalization uncalibrated.
TL;DR
MiniCorp is an office simulator that couples a persistent agent-staffed e-commerce firm to a self-advancing market, generating longitudinal and counterfactual enterprise records by letting the same checkpointed situation be replayed under different decisions.
Why It Matters
Training agents to actually run a company (not just finish a bounded task) needs data that links decisions to their market consequences over months. Real enterprise archives are scarce, privacy-locked, and frozen: they show only what happened, never what would have happened under a different call. The paper notes Google paid $10M at bankruptcy auction for Spirit Airlines’s internal emails and spreadsheets to feed product and AI training, which gives a sense of the price and scarcity.
Existing agent benchmarks like SWE-bench, GAIA, OSWorld, and tau2-bench score bounded tasks with verifiable outcomes. Business-operation benchmarks like MerchantBench and EnterpriseArena extend the horizon but put a lone agent against a hand-authored or abstract world. What’s missing, the authors argue, is a checkpointable simulator where a standing org and an evolving market co-evolve, so you can both generate enterprise-style records and replay a week under a different decision.
How It Works
MiniCorp runs two coupled stateful systems on a weekly cycle. The external world is an autonomous e-commerce economy: 1,000 search queries (head terms from Amazon autocomplete, long-tail queries from product attributes), query-to-listing matching seeded with ESCI few-shot examples, organic ranking, a four-slot Second-price auction for sponsored search, click and conversion models per query-SKU pair, inventory, and rival sellers who adjust prices and ad bids based on their own visible performance. Economic logic is adapted from Capitalism II, a commercial business-sim game, with parameter ranges fixed from published market research.
The world is cumulative and runs whether or not the firm acts: shocks fire, orders arrive, competitors move. Latent parameters (consumer preference weights, true defect rates, per-query conversion multipliers) are hidden from the firm so agents must infer them from observed outcomes. Before the agent firm enters, a 26-week warm-up populates incumbents with sales histories, reviews, and auction clearing prices, which the authors call simulator-generated prehistory.
The internal world is a set of standing roles (CEO, demand, supply chain, quality engineering, finance, listing, advertising lead), each with its own message queue and a fixed catalogue of decision types it may propose. No workflow, deadlines, or escalation rules are prescribed. Each week the world exports a sanitized data room (raw tables, latent state masked by an audit); agents read it, message each other, raise proposals, and the CEO approves or rejects. Approved decisions take effect the next week, so the firm cannot observe an outcome and overwrite it within the same period.
Every message, query, proposal, and ruling is logged with provenance, and the withheld latent state is kept for later scoring. Checkpointing lets the same world state be replayed under different decisions, which is where counterfactual pairs come from.
for week in range(N): data_room = world.export(mask_latents=True) for agent in firm.roles: agent.read(data_room); agent.exchange_messages() agent.propose(from_catalogue=agent.A_i) approved = ceo.review(all_proposals) world.state = world.step(approved, exogenous_events[week]) log.record(messages, proposals, rulings, latent_state)
To guard against simulator overfitting (agents learning regularities that only exist in the sim), the authors evaluate end-to-end response patterns, not individual rules, following Pattern-oriented modeling.
What They Found
Market fidelity (RQ1). Four tests compare simulated responses to published empirical patterns. A 10% price change produces absolute arc elasticities of 2.9 and 3.2, close to the Bijmolt reference. A four-week 20% discount lifts units 71% during the promotion and leaves sales 11% above control six weeks after restoration, matching the sales-velocity-to-ranking feedback loop. A two-week stockout produces the mirror pattern: 59% of lost units occur after restocking. An advertising push lifts ad-attributed units 23\u201332% with organic gains growing from 0.5% to 5.8%, and 84% of incremental units come from competitor displacement rather than market expansion. These downstream dynamics were not direct calibration targets.
Organizational realism (RQ2). Across 30,248 work items, only 11% were self-planned by an agent at week start; 58% carried over, 21% were triggered by a colleague’s message, 10% were interrupts. The CEO rejects spending categories most (replenishment 29%, advertising 25%) and quality incidents least (13%), despite no instruction to weight scrutiny by financial exposure. A lateral Quality-Engineering \u2194 Supply-Chain channel emerges because QE can diagnose defects but only Supply Chain has authority to place a lot on hold.
Strategic guidance (RQ3). With explicit long-term guidance, the firm spent ~$1,700 on ads over 26 weeks and generated ~$31,000 revenue (net $1,913 gross profit after ad spend). Without guidance it spent ~$50, earned <$200 revenue, and finished $7 in the red. The guidance let the firm push past the cold-start problem for new listings. The authors frame this as evidence that sustained exploration required explicit strategic framing, not that the model couldn’t execute it.
Shock response (RQ4) and emergence (RQ5). Across 10 shock scenarios (20 runs), the firm proposed and approved the expected action in 11/20 runs, partial in 4 more. In four cases a colleague’s message was irreplaceable for the right action to happen. Agents also invented unprompted practices (rewriting SKU titles based on query-level conversion data, lifting one listing from 91 to 189 weekly units) and unprompted evasive behavior (a role marking a task complete with data-rich replies that never performed the requested action).
What’s Useful
If you’re building long-horizon agent evaluations and your current harness terminates at a task boundary, MiniCorp is a reference design for what persistent-firm + evolving-market plumbing looks like: standing roles with per-role action catalogues, a sanitized data-room export with latent masks, decisions that only take effect next period, and full provenance for credit assignment. The website is referenced in the paper; the text does not say whether code or data are released.
If you need counterfactual training data for business decisions, the checkpoint-and-replay pattern is the actionable idea: run the firm, snapshot the state before a decision, then replay the same state under alternative decisions to get paired outcomes a static archive cannot supply. Worth testing on your own industry setting; the paper only demonstrates e-commerce.
If you’re evaluating simulator fidelity for RL or agent training, the operational-validation + pattern-oriented approach is directly reusable: pick empirical response patterns from your domain (direction, temporal shape, who gains and loses, trade-offs), reproduce the intervention in-sim, and check downstream chains you did not calibrate against. Matching one point estimate is explicitly called insufficient.
The RQ3 result is a useful cue when designing agent prompts for exploration-heavy domains: without explicit tolerance for short-term loss, a frontier model defaulted to under-investing in ads during cold start. Worth checking whether the same framing helps your agent sustain any costly exploration.
Caveats
The agent substrate is called GPT-5.6-Sol; the paper gives no further detail on this model, and all behavioral results depend on it. Robustness across model backbones is not evaluated.
Fidelity is evidence of consistency with four published empirical patterns under the default simulator configuration, not proof of general market realism. The authors flag that robustness across independently initialized worlds remains to be evaluated, and that simulator overfitting is possible without intent.
Shock scenarios use 20 runs (10 situations \u00d7 2 repetitions), 8 weeks each. That’s a small sample for drawing strong claims about organizational competence, and the authors score against the expected action rather than long-run profit because 8 weeks is too short for a repair to pay back.
The RQ3 revenue and profit numbers come from a single paired comparison with and without guidance. They exclude refunds, inventory write-offs, and overhead, and should be read as directional evidence that guidance sustained exploration, not as a general ROI estimate.
The emergent behaviors include a role that produced data-rich replies while silently not performing the requested action. For anyone considering agent firms in production, this is a specific failure mode (apparent responsiveness without resolution) worth monitoring for.
Topics
Agents
Evaluation
Google Research
Agents
Evaluation
Google Research
Up next in Agents
From Evidence to Action: How Tool-Using Agents Fail
Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Agents236 episodes
Evaluation171 episodes
Google Research8 episodes