Skip to content
santoshkandula.dev

← Writing

August 2026 · SEC-RAG-Eval

SEC-RAG-Eval: measuring RAG where it fails, not where it shines

A retrieval-augmented question-answering service over US SEC 10-K, 10-Q, and 8-K filings — and the evaluation harness that measures it. The service runs on GCP Cloud Run: FastAPI over Postgres pgvector, 15,192 chunks from 84 filings under text-embedding-3-large, answers generated by Claude Haiku 4.5, 118 tests, and CI that blocks retrieval regressions. The RAG system is the subject under test. The harness is the product.

The headline is recall@5 rising from 0.44 to 0.64 on FinanceBench-150, at $0.009 per query. Everything the harness caught on the way to that number is more useful than the number.

Recall is a bracket, not a point

Scoring the identical retrieval three ways gives three different answers: 0.09 under strict substring matching, 0.64 under the shipped fuzzy overlap, and 0.81 under embedding similarity. Same system, same run — a spread of roughly 8× that comes entirely from how a “hit” is defined. fuzzy(0.5) counts a hit when at least half of a question’s gold-evidence tokens appear in a retrieved chunk; strict matching wants the gold span verbatim; semantic matching asks whether the chunk is close in embedding space. The honest recall sits inside that bracket, and reporting 0.64 on its own would hide it.

The clean validation was the artifact

Recall is mostly a property of the matcher, not the retriever — so a 50-pair study asked which matcher a human actually agrees with. An early pass used a Claude Haiku proxy labeler and ranked the shipped fuzzy matcher best of three, at 0.67 agreement. Hand-labeling the same 50 pairs dropped that to 0.18, second of three, with no matcher beating chance at that sample size. The clean result was an artifact of the LLM labeler. The repo shipped it once, then caught it and published the correction.

How much of the win was real

The V0→V2 jump moved two variables at once — the embedding model and the chunk size — so the gain was decomposed on a full grid. About 80% of the headline is real system improvement, and about 20% is chunk size inflating the overlap metric: a larger chunk clears the 50%-token bar more easily, independent of retrieval quality. The arbiter is answer accuracy, which chunk size cannot inflate, and it climbs from 0.36 to 0.47 across the same change — almost all of it from the embedding model, not the chunk size.

Being right vs. declining to guess

Answer accuracy is 0.50 over all 150 questions, and about 0.74 of what the service attempts. The gap is the refusal rate: the grounded prompt declines when the evidence isn’t retrieved rather than guessing. Faithfulness — grounding, not correctness — scores 0.93 under an LLM judge that a second model audited and agreed with on 19 of 20 verdicts.

What I’d build next

The failure is localized. Dense retrieval degrades on tables, where tables@5 moved 0.32 to 0.70 but still trails prose. That points to table-aware parsing and a hybrid sparse-dense retriever — not a larger embedding model, which would lift the average and leave the actual gap intact.

Try it on the live demo, or read the code in the repository.