Agent Harness Benchmark: Codex vs Claude Code vs Hermes | HarnessRouter
HarnessRouter Benchmark 01
Same task. Eight configurations. About a 475× cost range.
HarnessRouter recorded the same controlled synthetic-data Care Prep task and input across Codex, Claude Code, and Hermes configurations. Measured cost ranged from 0.47 to 223 credits, while end-to-end latency ranged from 1m 25s to 4m 36s.
This is one controlled test, not a universal ranking. It is the first in a series of HarnessRouter benchmarks. We will publish more tasks, configurations, and evaluation methods over time.
View all results How we tested
Configurations
| Harness | Model | End-to-end latency | Cost |
|---|---|---|---|
| Hermes](/content/harnesses/hermes-agent/index.html) | gpt-5.2 | 2m 33s | 0.47 credits |
| Codex](/content/harnesses/codex/index.html) | gpt-5.2 | 4m 36s | 0.72 credits |
| Hermes](/content/harnesses/hermes-agent/index.html) | gpt-5.5 | 1m 25s | 40.7 credits |
| Codex](/content/harnesses/codex/index.html) | gpt-5.5 | 2m 46s | 59.7 credits |
| Hermes](/content/harnesses/hermes-agent/index.html) | claude-sonnet-4.6 | 2m 49s | 39.6 credits |
| Claude Code](/content/harnesses/claude-code/index.html) | claude-sonnet-4.6 | 1m 31s | 83.5 credits |
| Hermes](/content/harnesses/hermes-agent/index.html) | claude-opus-4.8 | 2m 42s | 150 credits |
| Claude Code](/content/harnesses/claude-code/index.html) | claude-opus-4.8 | 3m 18s | 223 credits |
Care Prep benchmark results for eight harness and model configurations
The approximately 475× figure compares 223 with 0.47 credits. Because the displayed measurements are rounded, the ratio is also approximate. The latency range is 276 ÷ 85, or approximately 3.2×. The lowest-cost configuration consumed 99.8% fewer credits than the highest-cost configuration.
Lowest measured cost Hermes + gpt-5.2
0.47 credits, compared with 223 credits for the highest-cost configuration in this test.
Lowest measured latency Hermes + gpt-5.5
1m 25s end to end, compared with 4m 36s for the slowest configuration in this test.
A visible tradeoff Claude Sonnet 4.6
Claude Code finished faster, while Hermes consumed fewer credits. The right choice depends on the workload objective.
Dominated configurations 6 of 8
For six of the eight configurations, another setup in this test was both cheaper and faster. Three out of four setups here had a strictly better alternative available.
The harness alone 1.5x to 2.1x
With the model held constant, switching only the harness moved cost by 1.5x to 2.1x and end-to-end latency by up to 1.95x in this test.
Where the extremes come from Model tier, then harness
Most of the 475x spread comes from model-tier economics; on the same model, the harness multiplies cost by up to 2.1x. Both choices are parameters you can route on.
HarnessRouter Care Prep benchmark capture, August 2026. Harness identifiers are truncated in the product interface. Open the full-size capture.
Observed history from the same controlled synthetic-data experiment
“Strict grounding pass” checks schema compliance, exact supplied facts, evidence-reference validity, and required human-review flag. It is not a composite quality score or clinical validation.
Test scope One controlled synthetic-data Care Prep experiment across eight harness and model configurations.
Observed history The experiment retained 5 observed runs per configuration and applied an objective validator to completed outputs.
Held constant The dataset, Task, Skill, schema, and output contract remained fixed across configurations.
Changed The selected agent harness and model configuration changed between rows.
Measurements The source capture records cost and end-to-end latency for each configuration. The observed history reports execution success, strict grounding pass rate, and p95 duration.
Quality boundary This publication does not combine quality into one score. The objective validator reports specific grounding checks; safety and workflow usefulness remain separate human-review criteria.
Interpretation These results apply to this task and configuration set. Results can change with the task, input, harness, model, environment, and configuration.
Your workload is the benchmark
Compare harnesses on the work your product actually runs.
HarnessRouter Cloud records execution traces across harness and model configurations so teams can compare cost, latency, and quality using their own production criteria.