Get Started
Home
Topics
Search
Library
7 min read · Agents · RAG · Added Oct 9 · Paper published Oct 7, 2026

RunningTab: Direct Workspace Interaction with Environment-Side Tabs

Source: research paper via Hugging Face Daily Papers
Long-horizon file agents forget numbers they literally read: baselines lose 51% of rubric points to extraction failures even with content in scrollback. RunningTab shifts bookkeeping to the harness, logging every read verbatim outside the context window, lifting Workspace-Bench pass rate from 31% to 47% without fine-tuning.
TL;DR
RunningTab makes the environment, not the model, keep a running record of what a workspace task owes, what files the agent has read, and what it listed but never opened, so content the agent already saw does not get crowded out of its context before it writes the deliverable.
Why It Matters
Imagine you point an LLM agent at a shared drive and ask it to produce a quarterly report that pulls numbers from a dozen spreadsheets and policy PDFs. The agent has terminal access: it can ls, grep, and cat any file. It opens the right annual report, the number appears in its scrollback, and then. after 40 more turns of listing folders and reading other files, it writes the report without that number. The paper measures this concretely: on a benchmark called Workspace-Bench, a baseline agent using plain Direct Corpus Interaction (DCI) loses 51.1% of its rubric points to extraction failures, and 19.8% of the specific numeric values it was literally shown end up missing from the final deliverable.
The closest prior approach is Direct Corpus Interaction (DCI), where the agent just runs shell commands against raw files instead of using a retriever and a vector index. That solves “how do I reach the files,” but it leaves tracking the task to the model’s context window, which is exactly where things fall off the end.
How It Works
The core move is to split responsibilities. The agent decides what the task is asking for. The environment (the harness that executes tool calls and feeds observations back) is responsible for remembering everything else, because it already sees every read and every directory listing anyway.
Concretely, RunningTab maintains a per-task tab with three parts:
•
Requirements: the agent, at the start, writes down one item per fact/figure/section the task asks for.
•
Read log: every time a tool call actually returns file content, the environment stores that excerpt verbatim along with the file path and the command that produced it. The model does not have to decide to save anything.
•
Candidates: every file path that appears in a listing but has not been opened is kept as a lead, ranked by BM25 against the task text.
The tab lives outside the context window, so old entries do not get displaced by new observations. The agent interacts with it through four operations: add (declare a requirement), list (get a review that pairs each open requirement with its best-matching read-log excerpts and best unopened candidate files), show (re-fetch a stored excerpt without re-reading the file), and resolve (close a requirement either as done with a citation to a read-log entry, or as unavailable with a reason). A resolve as done is rejected if the cited excerpt shares no terms with the requirement, so the agent cannot just claim completion. If the agent tries to stop with requirements still open, the environment sends back a finish check once, listing what is still open and which files it never opened.
for turn in range(max_turns): action = agent.step(context + status_line(tab)) obs = run_tool(action) if action.reads_file(): tab.read_log.append(Excerpt(obs, path, cmd)) for path in listed_paths(obs): tab.candidates.add(path, score=bm25(path, task)) if action.is_stop() and tab.has_open_requirements(): obs = finish_check(tab) # returned once context.append(obs)
Note what the agent writes to the tab across a whole task: just the requirements at the start and their resolutions at the end. Everything else, the excerpts, the candidate list, the review, the finish check, is produced by the environment.
What They Found
The authors evaluate on three benchmarks (Workspace-Bench, TheAgentCompany, and OfficeQA Pro) with three LLMs (GPT-5.4-Nano, DeepSeek-V4-flash, and Gemini 3.8 Flash), averaged over three runs. RunningTab wins on every benchmark, every metric, every model. The comparison that matters is against baselines where the model keeps the same kind of record: a TODO list at the start (Plan-and-Solve prompting-style), Self-Refine at the end, and a combined “Self-Tracking” that does both. Those model-kept baselines move the needle only a little; RunningTab moves it substantially more. On Workspace-Bench with Gemini 3.8 Flash, for example, the Rubric Pass Rate goes from 31.4% (plain DCI) to 46.9% (RunningTab), while Self-Tracking actually drops to 26.9%.
Two diagnostics support the mechanism claim. First, when a numeric value the rubric expects is shown to the agent but missing from its deliverable, the tab’s read log still holds 94.8% of those values. When the same LLM is given just the review (requirement plus best-matching excerpts), it recovers 45.8% of those masked values; given the whole read log, 78.1%. So the content is captured, and the review surfaces a large chunk of it with a small fraction of the text.
Second, the ablation isolates which piece matters. Removing “environmental capture” (so the environment no longer records the read log and candidates on its own, leaving only the requirements the agent wrote) drops Pass Rate back to 41.4%, essentially DCI’s level. Removing the review is the next biggest hit. The authors read this as: the gain comes mostly from the environment recording things the model didn’t choose to save, not from prompting the model to be more organized. A separate split by reading volume shows the gap over DCI roughly doubles on the reading-heaviest third of tasks versus the lightest, consistent with “this helps most when there is more to forget.”
One nuance worth flagging: different models use the tab differently. Gemini leans on the review, GPT-5.4 nano barely engages until the finish check prompts it (then edits the deliverable 17.1% of the time), and DeepSeek uses the candidate list to open more files. The authors present this as evidence the tab is a general scaffold rather than one that only clicks with a particular prompting style.
What’s Useful
If you are building a long-horizon agent that assembles a deliverable from many files (report generation, data room analysis, policy compilation), the main lesson is architectural: do not rely on the model to remember what it has seen across a long trajectory with context truncation. Have the harness capture file reads and listings verbatim, keyed by task, and surface them back on demand. This does not require fine-tuning or model access; it is a harness change.
A few concrete things worth testing in your own setup:
•
Store read observations verbatim with provenance (path + command), rather than letting a summarization step compress them. The paper’s point about Shown-but-Missing content assumes the raw excerpt is still retrievable.
•
Pair each declared requirement with its best-matching stored excerpt using cheap lexical scoring (BM25) before the agent writes. The review brought back more than half of what the full log provided while using much less text, which matters for your prompt budget.
•
Add a finish-check gate that returns open items once. In this paper it mattered more for some models than others, but on GPT-5.4 nano a majority of attempts touched the tab again after the check.
Things the paper does not establish, and you shouldn’t assume: that this helps for short-answer QA with small corpora (the OfficeQA Pro gains are smaller and that benchmark is still document-grounded, not single-shot), that it generalizes beyond file-system workspaces to, say, web navigation, or that any particular model will use the tab the way Gemini does rather than the way GPT-5.4 nano does. The authors do not report a public code release in the text supplied.
Caveats
The benchmarks are all file-based knowledge-work settings; the mechanism is tied to tool calls whose outputs the environment can classify as “read content” vs “listed path.” Ports to other action spaces are speculative. The tab stores raw file excerpts outside the context, which the authors note raises an access-control question for sensitive corpora: it should inherit the workspace’s permissions. The read log explicitly does not capture values the agent computed (sums, derivations) or values that appeared only in a listing, so the 94.8% recovery figure applies to numbers that showed up inside file contents. And while the ablation argues environmental capture is the load-bearing part, the finish check and evidence-gated resolve still contribute, so stripping RunningTab down to “just log everything” would likely give back some of the gain.
Topics
Agents
RAG
Agents
RAG
Up next in Agents
nanoMuse: An Open-Source Personal Agent for Every Device You Own
Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Agents244 episodes
RAG44 episodes