Reviewed comparison · August 12, 2026

Luna is the value leader. Keep Gemini 3.6 for detection.

The latest generally available models were tested against the complete six-case corpus: three layouts, each in clean and scanned form. Gemini candidates received three runs per case; every GPT-5.6 model completed one full MCP workflow per case. Production configuration was not changed.

6Clean and scanned cases
54Total paid API workflows
100%Field F1 and positioning · 3.6
$13.79Luna estimate / 1,000 workflows

Recommended stack

The cheapest agent matched the larger GPT-5.6 models, while the cheaper Gemini detector did not preserve field quality.

Best measured value
Provisional agent + retained detector

GPT-5.6 Luna + Gemini 3.6 Flash

Luna chose the correct MCP tools and preserved exact IDs in all six cases. Its sole strict failure was a shared PDFSight AcroForm accounting bug that affected Terra and Sol identically.

6 / 6Workflows executed
21.57sMean workflow latency
$0.01379Mean total workflow cost
AgentStrict$/flow
GPT-5.6 Luna5 / 6$0.01379
GPT-5.6 Terra5 / 6$0.02701
GPT-5.6 Sol5 / 6$0.05047
Correct tool + ID behavior6 / 6all
Infrastructure success100%all
Luna is 49% cheaper than Terra

On the measured mix, Luna costs $13.79 per 1,000 workflows versus $27.01 for Terra, with no measured quality loss.

Promotion still needs one fix and repeats

Repair the native AcroForm missing-count contract, then run three workflows per case. One repetition per case is directional, not a production confidence bound.

$0.01379
$0.02701
$0.05047

Latest Gemini detector comparison

Each model ran 18 times: six cases × three repeats. Strict success requires field discovery, type classification, and PDF positioning to pass together.

Keep incumbent
Production · keep

Gemini 3.6 Flash

The accuracy baseline. Perfect mean field F1 and positioning, with two type-threshold misses on the bookkeeper form.

Strict success16 / 18
Field / position100%
Type accuracy98.41%
p95 latency13.02s
Mean cost$0.01240
Reject

Gemini 3.5 Flash-Lite

Fast and inexpensive, but it loses fields, misclassifies types, fails positioning, and produced malformed JSON once.

Strict success7 / 17
Execution success94.44%
Field F169.63%
Type accuracy73.95%
Position pass76.47%
p95 / mean cost3.42s / $0.00363
$0.01240
$0.00363
13.02s
3.42s

Pricing options

Observed mean token usage, multiplied to planning volumes. USD standard synchronous API rates; taxes, negotiated discounts, Batch, Flex, Priority, and tool fees are excluded.

Value

Luna + Gemini 3.6

Best measured price/performance. Detector accounts for about 89% of the total.

Per workflow$0.01379
1,000 workflows$13.79
10,000 workflows$137.89
Balanced, no measured gain

Terra + Gemini 3.6

Nearly the same p95 latency as Luna, at roughly twice the total cost.

Per workflow$0.02701
1,000 workflows$27.01
10,000 workflows$270.14
Premium, no measured gain

Sol + Gemini 3.6

The most expensive and slowest option on this tool-routing workload.

Per workflow$0.05047
1,000 workflows$50.47
10,000 workflows$504.72
Detector-only planning

Gemini 3.6 averaged $0.01240 per PDF: about $12.40 per 1,000 or $123.96 per 10,000 on this mix.

The tempting cheap detector is unsafe

Gemini 3.5 Flash-Lite projects to $3.63 per 1,000, but its 41.18% completed-run success rate makes that saving operationally false economy.

Decision and limitations

The evidence supports a direction, but not an automatic production switch today.

Provisional choice: GPT-5.6 Luna

It matched Terra and Sol case-for-case, had the lowest mean latency, and was 49.0% cheaper than Terra and 72.7% cheaper than Sol.

Keep Gemini 3.6 Flash

Flash-Lite was 73.8% faster and 70.7% cheaper, but materially failed recall, type, positioning, and JSON reliability.

Fix the shared AcroForm bug

The application returns seven fields filled and seven missing for the same native form. That defect is why every GPT-5.6 model scored 5/6.

Expand confidence before promotion

Repeat agent runs three times per case and add rotated, handwritten, multilingual, mobile-capture, and independent real-world PDFs.

Source: evals/baselines/2026-08-12-latest-model-full-comparison.json Detector: 3 repetitions/case · Agent: 1 repetition/case · exact-model USD pricing