OCCAM: when is an agent worth its tokens?

RAG vs GraphRAG vs Agentic GraphRAG vs cost-aware routing, over 100 questions on the TigerGraph Agentic GraphRAG corpus.

Every number on this page is computed from the raw run files in artifacts/ at build time.

RAG
63%
3,259 tokens / question
GraphRAG
97%
838 tokens / question
Agentic GraphRAG
100%
890 tokens / question
OCCAM (routed)
100%
64 tokens / question

Cost against accuracy

Up and to the left is better. The agent reaches the top; routing reaches the same height for roughly 13× fewer tokens.

RAGGraphRAGAgentic GraphRAGOCCAM (routed)
0%25%50%75%100%1001,000mean tokens per question (log scale)answer accuracyRAG63% · 3,259 tokGraphRAG97% · 838 tokAgentic GraphRAG100% · 890 tokOCCAM (routed)100% · 64 tok

Where each approach breaks

RAG does not degrade gracefully on aggregation and superlative questions. Those need a median of 11 documents and up to 43, so the evidence does not fit in a top-10 retrieval whatever the prompt says.

RAGGraphRAGAgentic GraphRAGOCCAM (routed)
0%25%50%75%100%lookupn=19multi_hopn=28temporaln=22aggregationn=2110%superlativen=10

Completeness: did it investigate, or just retrieve?

Solid bars are answer accuracy; pale bars are completeness, the share of the question's gold documents the pipeline actually retrieved. A wide gap means answers are being produced without the evidence that settles them.

0%50%100%RAG63%76%GraphRAG97%98%Agentic GraphRAG100%99%OCCAM (routed)100%99%

Full metrics

pipelineaccuracycompletenessevidence precisiontokens / questiontokens / correct answerstepssec / q
RAG63%76%31%3,2595,1742.02.5
GraphRAG97%98%98%8388642.02.1
Agentic GraphRAG100%99%99%8908902.12.4
OCCAM (routed)100%99%99%64642.10.2

Where the agent tier earned its cost

Comparing the agentic run with single-pass GraphRAG on the same 100 questions:

questions the agent rescued 3 (3%)
questions the agent broke 0
extra tokens spent in total5,221
extra tokens per question rescued1,740
Rescued: pub-040, pub-055, pub-099. All three were cases where the graph narrowed the answer to a handful of candidates but could not choose between them. That is the shape of question worth paying an agent for.

How OCCAM routed

Each question starts at the cheapest tier and climbs only when a tier reports what it was missing.

tierquestions accuracytokens / question
tier0_rule_graph98100%0
tier1_disambiguated1100%263
tier3_document_fallback1100%6,161

Held-out set (50 questions)

Ground truth for these is not published, so only coverage and cost can be reported here - accuracy is scored by the organisers. Raw answers and full agentic traces are in artifacts/results_hidden.jsonl.

pipelineanswered tokens / questionsteps
RAG45/503,3322.0
GraphRAG47/508352.0
Agentic GraphRAG50/509522.2
OCCAM (routed)50/501422.3
OCCAM routing: tier0 rule graph ×48, tier3 document fallback ×1, tier1 disambiguated ×1