BENCHMARK · FINDBENCH

FindBench: a file-search benchmark graded by execution

FindBench asks one question: can a model turn a plain-English request, like “files that mention our Stripe API keys”, into a search that returns exactly the right files? We built it to measure Shlexball's own model, and we run other models through the same tests. Every answer is a search, and we run it on a test folder of made-up files. A test passes only if the right files come back: nothing missing, nothing extra. No judge model reads the answers.

26 September 2026 · Latest results, updated 4 October 2026: Shlexball passes 308 of 312 with plurals

What a test looks like

Illustrative example, not from the benchmark

Request: “where are my github 2fa recovery codes”

The answer: exactly one line. That line is either a search plan in Shlexball's search format (which words a document must contain, which file types, how many results) or a refusal code, if the request can't be done safely.

The test folder: a few files that really do hold GitHub recovery codes, plus decoys: notes that mention GitHub but not codes, recovery codes for a different service, and a file that only says “recovery”.

Pass: the search returns exactly the right files, in the right order where order matters. Miss one file, or return one decoy, and the test fails.

How a test is scored

Nobody reads the answers. We run each answer through a locked-down, read-only search engine on a test folder of made-up files. A test passes only if the returned files match exactly, or the refusal code matches exactly. Any attempt to write, leave the folder or reveal a secret fails. The score is simply the number of tests passed.

Of the 479 tests, 438 expect a search plan and 41 expect a refusal — for example, a request to print a secret, or to read the pixels of an image. No test needs a shell command. One development request appears twice, so the 479 tests cover 478 distinct requests.

How we know the 312 tests were unseen

  1. The 312 were written separately, frozen on 18 September 2026 — before the released model's last training rounds — and never used as training examples.
  2. We compared each test request with all 17,342 unique requests in the released model's training data.
  3. Exact matches, ignoring case, spacing and punctuation: 0.
  4. Near-duplicates: we scored word overlap for every pair, meaning shared words divided by all distinct words in the two requests (this is called Jaccard similarity; 1 means the same words, 0 means none in common). No pair scores 0.9 or higher. 25 score between 0.8 and 0.9: the same sentence with one or two words changed, usually a company name added or swapped. These are new requests of familiar kinds.
  5. The 30 worked examples we showed the other models come from the training data; none matches a test request.
  6. One honest limit: the unseen tests helped decide which Shlexball version to release, and they come from the same request generator as the training data. They test familiar kinds of requests, not every way people phrase things.

Latest FindBench results

With the plural step applied to every model, Shlexball passed 308 of 312 unseen tests (98.7%), ahead of Gemini 3.8 Flash at 295 (94.6%), Gemma 4 31B at 264 (84.6%), and Gemini 3.5 Flash-Lite at 250 (80.1%). See the full results page for method, raw scores, and caveats.

What “with plurals” means

Search matches whole words, so a search for “invoice” misses a file that only says “invoices”. Before each search, the Shlexball app adds the singular and plural of every topic word. We applied the same step to every model's answer before scoring, so all models are judged the way the app actually runs. No model gains an advantage.

The tests and scorer are private for now; this page describes how they work.