Genesis PreviewE1 admittedSemantic Snapshot implemented500 Atlas identitiesQuery Lab: liveRead-onlyNo arbitrary URLsHosted CI: startup blocked

Live read-only Genesis baseline plus protected E4.7 agent, ontology and model candidates. Public MCP, frontier-model execution and TNOE training are not deployed; no arbitrary URLs.

TWIRX Reference implementation and public service of the Typed Web Commons

AgentBench

Measure time to a verifiable answer—not just fluent output.

TWIRX compares the same agent under different Web-access conditions and publishes failures, abstentions, proof gaps, latency, tokens, network work and cost. The 100-scenario SOTA corpus is ready; its paid model conditions have not run.

Does TWIRX improve the same frontier agent?

The benchmark changes the agent’s Web-access condition, not the model. This isolates the value of reusable semantic state from general model capability.

AgentBench-SOTA conditions
ConditionAccessPurpose
Web searchSearch toolBroad current discovery baseline
BrowserControlled browserHuman-interface execution baseline
Native APIPublisher-specific APIBest known structured-source baseline
TWIRX MCPNine semantic toolsCompiled state, task-ready context and proof
TWIRX-first hybridTWIRX, then search for gapsCoverage and efficiency hypothesis
Deterministic ceilingGold typed query, no LLMCorpus and execution correctness

What has actually run

E4.6: completed

  • 40 gold scenarios
  • 100% deterministic correctness
  • 100% proof completeness
  • 0.036 ms deterministic p95

E4.6: published loss

  • HuggingFaceTB/SmolLM2-360M-Instruct
  • 0% valid success under the strict answer contract
  • No product benefit attributed to the small planner
  • No inference about a frontier model

The deterministic bounded system passed; the small conversational planner was not suitable for the strict answer contract. This does not test a frontier model.

New local binding measurement

For one four-field World State query, the default twirx_context response used 9,517 encoded bytes, versus 13,089 bytes for separate query, compact-trace and explain responses—a measured 27.29% reduction. This measures the local response representation, not tokens, answer quality, latency or provider cost.

What is ready but not executed

Gold scenarios
100
Unanswerable
26
Planned model runs
5,000
Executed model runs
0

Opportunity and World State each contribute 50 scenarios. Every stochastic condition repeats 10 times. The same provider model, prompt policy, structured answer schema and spend bounds apply across model-driven conditions.

No decorative benchmark

The corpus and execution plan are complete. Model results remain zero until a founder-approved cost pilot runs. Recorded replays and live calls will never be combined under one label.

Primary endpoint: Time to Verifiable Answer

Elapsed time stops only when each substantive claim is either evidence-bound or withheld.

Answer quality

Correctness, numeric and temporal accuracy, constraint satisfaction, abstention.

Evidence

Claim-level proof coverage, invalid references, native-value recoverability.

Efficiency

Latency, tokens, tool calls, requests, bytes, browser launches and cost.

Reliability

Unsupported claims, stale-data errors, source disagreement and unknown-state preservation.

Stability

Answer, tool-plan, citation, failure and cost variance across repeated runs.

System cost

Model spend, deterministic compute, proof bytes and operational work are reported separately.

TWIRX may lose

  • Search may win on newly published or uncompiled information.
  • A native API may win for one known source and one narrow call.
  • A browser may remain necessary for an unknown interactive origin.
  • TWIRX is expected to be strongest on repeated, structured, multi-source and proof-sensitive work—but AgentBench must test rather than assume that claim.
  • Every timeout, refusal, schema error, unsupported claim and cost overrun remains in the release evidence.

Inspect the agent interface Read the benchmark report