BENCHMARK · FINDBENCH
FindBench asks one question: can a model turn a plain-English request, like “files that mention our Stripe API keys”, into a search that returns exactly the right files? We built it to measure Shlexball's own model, and we run other models through the same tests. Every answer is a search, and we run it on a test folder of made-up files. A test passes only if the right files come back: nothing missing, nothing extra. No judge model reads the answers.
26 September 2026 · Latest results, updated 4 October 2026: Shlexball passes 308 of 312 with plurals
Illustrative example, not from the benchmark
Request: “where are my github 2fa recovery codes”
The answer: exactly one line. That line is either a search plan in Shlexball's search format (which words a document must contain, which file types, how many results) or a refusal code, if the request can't be done safely.
The test folder: a few files that really do hold GitHub recovery codes, plus decoys: notes that mention GitHub but not codes, recovery codes for a different service, and a file that only says “recovery”.
Pass: the search returns exactly the right files, in the right order where order matters. Miss one file, or return one decoy, and the test fails.
Nobody reads the answers. We run each answer through a locked-down, read-only search engine on a test folder of made-up files. A test passes only if the returned files match exactly, or the refusal code matches exactly. Any attempt to write, leave the folder or reveal a secret fails. The score is simply the number of tests passed.
Of the 479 tests, 438 expect a search plan and 41 expect a refusal — for example, a request to print a secret, or to read the pixels of an image. No test needs a shell command. One development request appears twice, so the 479 tests cover 478 distinct requests.
With the plural step applied to every model, Shlexball passed 308 of 312 unseen tests (98.7%), ahead of Gemini 3.8 Flash at 295 (94.6%), Gemma 4 31B at 264 (84.6%), and Gemini 3.5 Flash-Lite at 250 (80.1%). See the full results page for method, raw scores, and caveats.
Search matches whole words, so a search for “invoice” misses a file that only says “invoices”. Before each search, the Shlexball app adds the singular and plural of every topic word. We applied the same step to every model's answer before scoring, so all models are judged the way the app actually runs. No model gains an advantage.
The tests and scorer are private for now; this page describes how they work.