Project 05 · Retrieval · 2026

Does grounding still beat a frontier model in Korean?

In progress · Phase 3

A benchmark comparing grounded retrieval against frontier models answering from parameters alone, over Korean corporate filings from DART. The English-language version of this question is close to settled. The Korean one is not, and the gap is where the interesting answer lives.

[TODO: one paragraph on method — multilingual-E5 for embeddings, dense-only retrieval, and the Mr. TYDI Korean Wikipedia notebooks that justify both choices. Say plainly that those two decisions were made and frozen before the DART results existed.]

[TODO: what you expect to find, written now, before you have the numbers. A pre-registered guess you can be wrong about in public is worth more than a clean result explained afterwards.]

RoleSole author
CorpusDART corporate filings
BaselinesGPT-4o, Claude, Gemini
StatusTest set in construction