Lesson 4 · Comparable measurementCourse indexPT-BR
CLIProxyAPI benchmark · Comparable measurement

Deliver the context you claim

The nominal target changes when the cap, envelope, and output reserve share one window.

Primary local sources: BENCHMARK-MEASUREMENT-PLAN.md, model-context-matrix.csv, and planned-run-manifest.json.
BENCHMARK-MEASUREMENT-PLAN.md · model-context-matrix.csv · planned-run-manifest.json

The idea in plain language

The matrix uses five scenarios: 256, 128k, 256k, 500k, and ~1M. ✓ fits a known limit; △ needs proof; — exceeds it; N/A is non-textual. A harness cannot claim it tested 500k if it compacted or truncated before transmission.

Running example: send the same sealed parcel — nonce, checksum, and three needles — through all six lanes. If the harness changes the wrapping, you are measuring the lane too. The analogy breaks because context, tools, and thinking carry semantics, not size alone.

The actual target is min(nominal scenario, effective cap − observed envelope − output reserve). A full lane has 53 strict and 9 conditional cells. If all six lanes qualified, the geometric upper bound would be 1,278 requests and 238,024,602 known input tokens — a projection, not authorization.

Verifiable excerpt

actual_target = min(nominal, effective_cap - envelope - output_reserve)\nper_lane = 53 strict + 9 conditional\nall_six_lanes_upper_bound = 1278 requests

outputs/BENCHMARK-MEASUREMENT-PLAN.md · outputs/model-context-matrix.csv · outputs/planned-run-manifest.json

The path in one picture

UI / CLIobserved outputharnessWarp · Codex · Claudecompatibilitysidecar / envelopeCLIProxyAPI127.0.0.1:8317providerdirect Chat · Responses · MessagesPRIMARY BOUNDARY
Running example: the same nonce + checksum crosses every lane; the failure belongs to the smallest boundary where it reproduces.

Retrieval check

Can hidden compaction compare with the direct grid?