The September 2026 Price Ladder: Building a Cost-Aware Model Router

Opus 5.5 and Sonnet 5.5 landed in September 2026. Verified Anthropic prices, the tokenizer trap, and a tested TypeScript router that cascades cheap to strong.

By Ajith joseph · Tue Sep 29 2026 · Updated Tue Sep 29 2026 · 11 min read · intermediate

#typescript #ai #claude #llm #cost-optimization

September 2026 was a busy month for model releases, and for once the interesting part is not the benchmark table. It is the price list.

Anthropic shipped Claude Opus 5.5 on 22 September and Claude Sonnet 5.5 on 28 September. Other labs shipped too, and by some trackers' counts there were two dozen releases across a couple of weeks. The practical question for anyone running an AI feature is not which model is best. It is which model is good enough for each request, at what price, and how to make that decision in code instead of in a meeting.

This post takes the prices Anthropic publishes, works out what they mean for real request costs, points out a trap in comparing them, and builds a small router in TypeScript with a test that checks the arithmetic.

What Anthropic Announced

The two announcements, as worded on Anthropic's newsroom:

  • Opus 5.5 (22 September): it "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5."
  • Sonnet 5.5 (28 September): "a clear upgrade over Sonnet 5 that runs 30% faster and costs up to 30% less for most work."

Read those claims carefully. They are about cost to run, meaning the price of finishing a piece of work, not the price per token. Those are different numbers, and the difference is the useful part of this post.

The Price List

These are the per-million-token prices from Anthropic's pricing page as of 29 September 2026. Confirm them before you rely on them, because prices change and I have copied only what the page showed.

Model Input Output Cache read 5-minute cache write
Claude Haiku 4.5 $1 $5 $0.10 $1.25
Claude Sonnet 5.5 $2 $10 $0.20 $2.50
Claude Opus 5.5 $4 $20 $0.20 $5
Claude Opus 5 $5 $25 $0.50 $6.25
Claude Fable 5.1 $10 $50 $0.25 $12.50

Some details in the fine print change the maths:

  • Opus 5.5's list price is 20% below Opus 5 ($4 and $20 against $5 and $25), but Anthropic claims 40% lower cost to run. The gap most likely comes from doing the same work with fewer tokens, which depends on your workload. Anthropic's page does not break it down, so treat 40% as a claim to test on your own prompts, not a number to budget with.
  • Cache reads are cheaper on the new models. Opus 5.5 cache reads cost 5% of the input price rather than the usual 10%, which is why its cache read is $0.20 at a $4 input price. Fable 5.1 is lower again at 2.5%.
  • Sonnet 5's introductory $2 and $10 became permanent. The scheduled rise to $3 and $15 on 1 September will not happen, according to a footnote on the page.
  • Batch is half price. The Batch API takes 50% off input and output, and it stacks with caching.
  • Fast mode is a premium. For Opus 5.5 it is $8 input and $40 output, and it is not available with batch.
  • US-only inference costs 1.1 times as much on the first-party API for Claude 4.6 and later.

The Tokenizer Trap

This one changes how you compare models, and it is easy to miss. Anthropic's page says Claude 4.7 and later models use a newer tokenizer that "produces approximately 30% more tokens for the same text," depending on content. Claude Sonnet 4.6 and earlier use the previous tokenizer.

Haiku 4.5 is on the older tokenizer. So the same prompt is roughly 30% fewer tokens on Haiku than on Sonnet 5.5. A $1 per million price is therefore not a clean 2 times cheaper than Sonnet's $2. Per request, Haiku is cheaper than the list prices suggest, and comparing per-token prices across the boundary understates that.

This is also part of why "cost per token" and "cost to run" diverge on newer models: the token counts themselves are not comparable across generations.

What a Request Costs

Take a workload of 3,000 input tokens and 500 output tokens per request, counted on the modern tokenizer, with the Haiku figure scaled for the older one. This is my arithmetic from the published prices, and the test in the next section checks it:

Model Per request Batch (half price)
Haiku 4.5 $0.00423 $0.00212
Sonnet 5.5 $0.01100 $0.00550
Opus 5.5 $0.02200 $0.01100
Fable 5.1 $0.05500 $0.02750

Sonnet 5.5 is 3,000 times $2 plus 500 times $10, divided by a million: $0.011. Opus 5.5 doubles it. Fable 5.1 is five times Sonnet.

With prompt caching, a request where 2,500 of those 3,000 input tokens are cache reads on Sonnet 5.5 costs $0.0065 instead of $0.011, a 41% saving from caching alone. Anything with a long, stable system prompt should be cached before you start routing, because it is the cheaper optimisation.

The Router

Two ideas are worth encoding. First, cost is a function you can compute exactly, so compute it. Second, a cascade tries the cheap model first and only escalates when a check rejects the answer.

// Prices are USD per million tokens, from
// https://platform.claude.com/docs/en/about-claude/pricing on 2026-09-29.
export interface Price {
  input: number;
  output: number;
  cacheRead: number;
  cacheWrite5m: number;
  // Claude 4.7+ tokenizers emit ~30% more tokens than Sonnet 4.6 and earlier for the
  // same text. Haiku 4.5 uses the older tokenizer, so a prompt is ~1/1.3 the tokens.
  tokenScale: number;
}

export const PRICES = {
  'claude-haiku-4-5-20251001': { input: 1, output: 5, cacheRead: 0.1, cacheWrite5m: 1.25, tokenScale: 1 / 1.3 },
  'claude-sonnet-5-5': { input: 2, output: 10, cacheRead: 0.2, cacheWrite5m: 2.5, tokenScale: 1 },
  'claude-opus-5-5': { input: 4, output: 20, cacheRead: 0.2, cacheWrite5m: 5, tokenScale: 1 },
  'claude-fable-5-1': { input: 10, output: 50, cacheRead: 0.25, cacheWrite5m: 12.5, tokenScale: 1 },
} as const satisfies Record<string, Price>;

export type ModelId = keyof typeof PRICES;

export interface Usage {
  inputTokens: number; // uncached input, counted with the modern tokenizer
  cacheReadTokens?: number;
  cacheWriteTokens?: number;
  outputTokens: number;
}

/** Dollar cost of one call. `batch` applies the 50% Batch API discount to every line. */
export function cost(model: ModelId, u: Usage, batch = false): number {
  const p: Price = PRICES[model];
  const s = p.tokenScale;
  const dollars =
    (u.inputTokens * s * p.input +
      (u.cacheReadTokens ?? 0) * s * p.cacheRead +
      (u.cacheWriteTokens ?? 0) * s * p.cacheWrite5m +
      u.outputTokens * s * p.output) /
    1_000_000;
  return batch ? dollars * 0.5 : dollars;
}

/** Cheapest tier first; escalate only when `accept` rejects the answer. */
export async function cascade<T>(
  tiers: readonly ModelId[],
  attempt: (model: ModelId) => Promise<{ result: T; usage: Usage }>,
  accept: (result: T) => boolean,
): Promise<{ result: T; model: ModelId; spent: number }> {
  let spent = 0;
  for (let i = 0; i < tiers.length; i++) {
    const model = tiers[i];
    const { result, usage } = await attempt(model);
    spent += cost(model, usage);
    if (accept(result) || i === tiers.length - 1) return { result, model, spent };
  }
  throw new Error('cascade needs at least one tier');
}

/** Expected per-request cost when the cheap tier is accepted a fraction `p` of the time. */
export function expectedCascadeCost(cheap: number, strong: number, p: number): number {
  return cheap + (1 - p) * strong;
}

/** Acceptance rate above which trying `cheap` first beats always calling `strong`. */
export function breakEvenAcceptance(cheap: number, strong: number): number {
  return cheap / strong;
}

I used claude-opus-5-5 and claude-sonnet-5-5 as the model IDs. They were not on the pricing page, so confirm the exact IDs in Anthropic's models list before shipping. The test hand-computes every figure above and runs the cascade, and it passes under Node's built-in type stripping:

import assert from 'node:assert/strict';
import { PRICES, breakEvenAcceptance, cascade, cost, expectedCascadeCost, type Usage } from './router.ts';

const close = (a: number, b: number) => assert.ok(Math.abs(a - b) < 1e-9, `${a} !== ${b}`);
const call: Usage = { inputTokens: 3000, outputTokens: 500 };

close(cost('claude-sonnet-5-5', call), 0.011);          // 3000*2 + 500*10 micro-dollars
close(cost('claude-opus-5-5', call), 0.022);
close(cost('claude-fable-5-1', call), 0.055);
close(cost('claude-sonnet-5-5', call, true), 0.0055);   // Batch API: half price

// 2,500 of the 3,000 input tokens served from cache on Sonnet 5.5
close(cost('claude-sonnet-5-5', { inputTokens: 500, cacheReadTokens: 2500, outputTokens: 500 }), 0.0065);

// Opus 5.5 cache reads are 5% of input, not the usual 10%
close(PRICES['claude-opus-5-5'].cacheRead / PRICES['claude-opus-5-5'].input, 0.05);

close(breakEvenAcceptance(cost('claude-sonnet-5-5', call), cost('claude-opus-5-5', call)), 0.5);
close(expectedCascadeCost(0.011, 0.022, 0.8), 0.011 + 0.2 * 0.022);

// The cascade escalates once on rejection and adds up what it spent...
const seen: string[] = [];
const out = await cascade(
  ['claude-sonnet-5-5', 'claude-opus-5-5'],
  async (model) => {
    seen.push(model);
    return { result: model, usage: call };
  },
  (r) => r === 'claude-opus-5-5',
);
assert.deepEqual(seen, ['claude-sonnet-5-5', 'claude-opus-5-5']);
close(out.spent, 0.011 + 0.022);

// ...and never touches the strong tier when the cheap answer is accepted.
const only: string[] = [];
await cascade(['claude-sonnet-5-5', 'claude-opus-5-5'], async (m) => (only.push(m), { result: 'ok', usage: call }), () => true);
assert.deepEqual(only, ['claude-sonnet-5-5']);

When a Cascade Pays

A cascade pays when the cheap model's answer is accepted often enough. The expected cost is the cheap call plus the strong call weighted by how often you escalate, and the break-even acceptance rate is simply the cheap price divided by the strong price.

Anthropic's list prices double at each step from Haiku to Sonnet 5.5 to Opus 5.5. So for adjacent tiers the break-even is 50%: if Sonnet 5.5's answer is accepted more than half the time, trying it before Opus 5.5 is cheaper than always calling Opus.

Because of the tokenizer effect, Haiku to Sonnet is better than that. Haiku's effective cost per request is $0.00423 against Sonnet's $0.011, so break-even is about 38%.

Worked example, with every assumption stated: 100,000 requests a month at 3,000 input and 500 output tokens, Sonnet 5.5 tried first with 80% acceptance, Opus 5.5 as the fallback, and a verifier that costs nothing.

  • Always Opus 5.5: 100,000 times $0.022 is $2,200
  • Always Sonnet 5.5: 100,000 times $0.011 is $1,100
  • Cascade at 80% acceptance: $0.011 plus 0.2 times $0.022 is $0.0154 per request, so $1,540, a 30% saving against always-Opus

Those numbers are only as good as the 80%, which you must measure, and the free verifier, which you will not get.

What This Model Leaves Out

Be suspicious of your own spreadsheet. Four things sit outside it:

  • The verifier is not free. Something has to decide whether an answer is acceptable. If that is another model call, add its cost to every request. Cheap deterministic checks, such as schema validation, a citation that must resolve, or a number that must reconcile, are the best kind.
  • Escalation adds latency. A rejected first attempt means a user waits for two calls. If your interface is interactive, that matters more than the saving.
  • Agents are not single calls. In a multi-step agent loop, a weak model that makes one wrong tool call can cost more than the strong model's whole run. Route by task type for agents, and keep confidence-based cascades for bounded jobs like classification, extraction and drafting.
  • Cost to run is not cost per token. Anthropic's own claim for Opus 5.5 is a 40% cost reduction on a 20% price cut. Measure tokens per completed task on your own prompts, because that ratio, not the price sheet, is your real cost.

I also limited the table to prices I could confirm on the vendor's own page. Trackers this month report very low prices from other providers, including a new floor around ten cents per million input tokens, but I could not load those vendors' pricing pages, so I have not put their numbers in a router. Add your other providers to the PRICES table with figures from their own pricing pages, and the rest of the code works unchanged.

A Practical Order of Operations

  1. Turn on prompt caching for any stable prefix. It is the largest saving with the least risk
  2. Use the Batch API for anything that can wait, since it is half price
  3. Build an evaluation set of real requests with a definition of acceptable, so acceptance rate is a measured number
  4. Route by task class first, with a static table from task type to model
  5. Add a cascade only to task classes with a cheap, reliable check
  6. Track cost per completed task, not per token, and re-run the comparison when a new model ships

New models will keep arriving on a monthly cadence, and each one moves the ladder. A router with the prices in one table, and a test that checks the arithmetic, turns that from a quarterly re-architecture into a config change.

Sources

  • Claude pricing, Anthropic, retrieved 29 September 2026
  • Anthropic newsroom, announcements for Claude Opus 5.5 (22 September 2026) and Claude Sonnet 5.5 (28 September 2026)
  1. AJ's Tech Notes
  2. The September 2026 Price Ladder: Building a Cost-Aware Model Router