BENCHMARK · FINDBENCH

FindBench results: Shlexball leads on 312 unseen file search tests

Shlexball leads FindBench, our file search benchmark of 312 unseen tests. With the plural step applied to every model, Shlexball passed 308 of 312 tests (98.7%). Gemini 3.8 Flash passed 295 (94.6%), Gemma 4 31B passed 264 (84.6%), and Gemini 3.5 Flash-Lite passed 250 (80.1%). The lead over Gemini 3.8 Flash is 13 tests. Shlexball's model has 1.5B parameters, about a twentieth of Gemma 4 31B's size.

26 September 2026 · updated 4 October 2026 at 11:47 am PDT (what changed)

Pixel-art illustration: a small gold robot beside a results board, above a grid of sealed test drawers with pass and fail lights.
Illustration, not to scale. The real results are in the board below.

The short version

Results

Every model received the same instructions Shlexball's app gives its own model. The comparison models also got 30 worked examples, and the board shows them with instructions only too. Shlexball's 1.5B model ran on a MacBook Pro (M4 Max). Gemini 3.8 Flash and Gemini 3.5 Flash-Lite ran on Google's servers, and Google does not disclose their sizes. Gemma 4 31B ran as the full uncompressed model on a rented NVIDIA H200 datacenter GPU.

What “with plurals” means

Search matches whole words, so a search for “invoice” misses a file that only says “invoices”. Before each search, the Shlexball app adds the singular and plural of every topic word. We applied the same step to every model's answer before scoring, so all models are judged the way the app actually runs. No model gains an advantage.

ModelParametersRuns onGivenWith plurals (312 unseen)
Shlexball (jeremiah v6)1.5BMacBook Pro (M4 Max)instructions308/312 (98.7%)
Gemini 3.8 Flashnot disclosedGoogle's serversinstructions + 30 examples295/312 (94.6%)
Gemma 4 31B31BNVIDIA H200 (RunPod)instructions + 30 examples264/312 (84.6%)
Gemini 3.5 Flash-Litenot disclosedGoogle's serversinstructions + 30 examples250/312 (80.1%)
Gemini 3.8 Flashnot disclosedGoogle's serversinstructions only203/312 (65.1%)
Gemma 4 31B31BNVIDIA H200 (RunPod)instructions only200/312 (64.1%)
Gemini 3.5 Flash-Litenot disclosedGoogle's serversinstructions only155/312 (49.7%)

Jeremiah v6 is the version name of Shlexball's model, fine-tuned from Alibaba's open Qwen2.5-Coder-1.5B-Instruct.

Updated 4 October 2026 at 11:47 am PDT. The main results now apply the plural step to every model, so all models are scored the way the app runs a search. Gemini 3.8 Flash was added to the board. No model was run again; we scored the answers each model had already given. Earlier the same day, at 11:37 am PDT, the plural step was first added as an extra column.

In the raw run, Shlexball produced 0 unsafe results and 0 wrong refusals.

Speed

Shlexball turned each request into a search in 0.41 seconds on average, faster than Gemini 3.5 Flash-Lite's 0.59 seconds — and Shlexball's number includes no network round trip, while Gemini's does.

0.41 sShlexball, on the Mac itself: average per answer, median 0.36 s
0.59 sGemini 3.5 Flash-Lite, Google's API: average per answer, median 0.59 s

How we measured: same 479 FindBench runs, on an M4 Max MacBook Pro (48 GB). Shlexball ran on the Mac itself, no network round trip; Gemini via Google's API, network included, 12 requests in flight. An “answer” means turning the request into a search, not searching your files. Gemma 4 31B is not in the speed comparison: we ran it on a datacenter GPU we rented and set up ourselves, with up to 64 requests at once, so its time would measure our setup, not a service you could use. Gemini 3.5 Flash-Lite is timed through Google's public API, which anyone can call. Slower Macs will be slower.

Every search starts with this step, so at well under half a second it never becomes the wait — the time you actually notice is the search over your files, which depends on how many you have.

How FindBench works

What a test looks like, how each answer is scored, what the 479 tests contain and how we checked that the 312 unseen tests never appeared in training are on the FindBench page.

Why a 1.5B model can win here

File search is a narrow job with a strict answer format and an exact right answer. With plurals, the app supplies singular and plural forms for every model, so the remaining gap comes from other conventions a model trained for this job learns: search for both “api_key” and “api-key”, and tell “a file that mentions a key” apart from “a file that stores a key”. A general-purpose model has to work these rules out from the instructions alone. Size helps when a task is broad. This task is not.

What the other models got

Limits

Benchmarks aside, the test that matters is your own files. Download Shlexball and try it: the first 30 searches are free.

▶ DOWNLOAD

macOS 15 or later. Apple silicon. 16 GB memory.