Does TWIRX improve the same frontier agent?
The benchmark changes the agent’s Web-access condition, not the model. This isolates the value of reusable semantic state from general model capability.
| Condition | Access | Purpose |
|---|---|---|
| Web search | Search tool | Broad current discovery baseline |
| Browser | Controlled browser | Human-interface execution baseline |
| Native API | Publisher-specific API | Best known structured-source baseline |
| TWIRX MCP | Nine semantic tools | Compiled state, task-ready context and proof |
| TWIRX-first hybrid | TWIRX, then search for gaps | Coverage and efficiency hypothesis |
| Deterministic ceiling | Gold typed query, no LLM | Corpus and execution correctness |
What has actually run
E4.6: completed
- 40 gold scenarios
- 100% deterministic correctness
- 100% proof completeness
- 0.036 ms deterministic p95
E4.6: published loss
HuggingFaceTB/SmolLM2-360M-Instruct- 0% valid success under the strict answer contract
- No product benefit attributed to the small planner
- No inference about a frontier model
The deterministic bounded system passed; the small conversational planner was not suitable for the strict answer contract. This does not test a frontier model.
New local binding measurement
For one four-field World State query, the default twirx_context
response used 9,517 encoded bytes,
versus 13,089 bytes for separate
query, compact-trace and explain responses—a measured
27.29% reduction. This measures the
local response representation, not tokens, answer quality, latency or provider cost.
What is ready but not executed
- Gold scenarios
- 100
- Unanswerable
- 26
- Planned model runs
- 5,000
- Executed model runs
- 0
Opportunity and World State each contribute 50 scenarios. Every stochastic condition repeats 10 times. The same provider model, prompt policy, structured answer schema and spend bounds apply across model-driven conditions.
No decorative benchmark
The corpus and execution plan are complete. Model results remain zero until a founder-approved cost pilot runs. Recorded replays and live calls will never be combined under one label.
Primary endpoint: Time to Verifiable Answer
Elapsed time stops only when each substantive claim is either evidence-bound or withheld.
Answer quality
Correctness, numeric and temporal accuracy, constraint satisfaction, abstention.
Evidence
Claim-level proof coverage, invalid references, native-value recoverability.
Efficiency
Latency, tokens, tool calls, requests, bytes, browser launches and cost.
Reliability
Unsupported claims, stale-data errors, source disagreement and unknown-state preservation.
Stability
Answer, tool-plan, citation, failure and cost variance across repeated runs.
System cost
Model spend, deterministic compute, proof bytes and operational work are reported separately.
TWIRX may lose
- Search may win on newly published or uncompiled information.
- A native API may win for one known source and one narrow call.
- A browser may remain necessary for an unknown interactive origin.
- TWIRX is expected to be strongest on repeated, structured, multi-source and proof-sensitive work—but AgentBench must test rather than assume that claim.
- Every timeout, refusal, schema error, unsupported claim and cost overrun remains in the release evidence.