Gemini 3.6 Flash
The accuracy baseline. Perfect mean field F1 and positioning, with two type-threshold misses on the bookkeeper form.
The latest generally available models were tested against the complete six-case corpus: three layouts, each in clean and scanned form. Gemini candidates received three runs per case; every GPT-5.6 model completed one full MCP workflow per case. Production configuration was not changed.
The cheapest agent matched the larger GPT-5.6 models, while the cheaper Gemini detector did not preserve field quality.
Luna chose the correct MCP tools and preserved exact IDs in all six cases. Its sole strict failure was a shared PDFSight AcroForm accounting bug that affected Terra and Sol identically.
On the measured mix, Luna costs $13.79 per 1,000 workflows versus $27.01 for Terra, with no measured quality loss.
Repair the native AcroForm missing-count contract, then run three workflows per case. One repetition per case is directional, not a production confidence bound.
Each model ran 18 times: six cases × three repeats. Strict success requires field discovery, type classification, and PDF positioning to pass together.
The accuracy baseline. Perfect mean field F1 and positioning, with two type-threshold misses on the bookkeeper form.
Fast and inexpensive, but it loses fields, misclassifies types, fails positioning, and produced malformed JSON once.
Observed mean token usage, multiplied to planning volumes. USD standard synchronous API rates; taxes, negotiated discounts, Batch, Flex, Priority, and tool fees are excluded.
Best measured price/performance. Detector accounts for about 89% of the total.
Nearly the same p95 latency as Luna, at roughly twice the total cost.
The most expensive and slowest option on this tool-routing workload.
Gemini 3.6 averaged $0.01240 per PDF: about $12.40 per 1,000 or $123.96 per 10,000 on this mix.
Gemini 3.5 Flash-Lite projects to $3.63 per 1,000, but its 41.18% completed-run success rate makes that saving operationally false economy.
The evidence supports a direction, but not an automatic production switch today.
It matched Terra and Sol case-for-case, had the lowest mean latency, and was 49.0% cheaper than Terra and 72.7% cheaper than Sol.
Flash-Lite was 73.8% faster and 70.7% cheaper, but materially failed recall, type, positioning, and JSON reliability.
The application returns seven fields filled and seven missing for the same native form. That defect is why every GPT-5.6 model scored 5/6.
Repeat agent runs three times per case and add rotated, handwritten, multilingual, mobile-capture, and independent real-world PDFs.
evals/baselines/2026-08-12-latest-model-full-comparison.json
Detector: 3 repetitions/case · Agent: 1 repetition/case · exact-model USD pricing