Pricing

Try it without an account. Every plan reduces context by the same amount; paid plans raise how often you can do it, and add the tuned adapter and team features.

Free

100 retrievals a month

$0 · no account required

  • Full context reduction
  • Every context window
  • Local search and context packets
  • The measurement dashboard
  • Community support on GitHub
Install for VS Code

Pro

300 retrievals a month

$29 USD a month

  • Everything in Free
  • The tuned adapter
  • Synced lesson store
  • Usage dashboard
  • Quarterly adapter updates
  • Email support, next business day
Choose Pro

Pro+

Unlimited retrievals

$89 USD a month

  • Everything in Pro
  • No monthly retrieval limit
  • Priority email support, next business day
Choose Pro+

Team

$69 per seat a month

5 seat minimum · unlimited retrievals · $59 a seat on annual billing, by email

  • Everything in Pro+
  • Shared lesson store
  • Team dashboard
  • Policy control over escalation
Choose Team

Enterprise

From $3,500 a month

plus $10,000–15,000 setup · unlimited retrievals

  • Self-hosted second tier
  • Custom adapter
  • Audit log export
  • SSO
  • Support SLA
Talk to me

Pro, Pro+ and Team check out through Stripe. Enterprise is arranged by email, because each deployment is built by hand.

Latorium is a software product operated and sold by Christian Yip Jun, an individual seller. Stripe securely processes subscription payments. See the Terms before purchasing.

What the limit means

The limit is a monthly count of retrievals. It is not a cap on how much context Latorium removes from any one of them.

A retrieval is one call from your agent — a context packet or a local search. Every plan reduces context by exactly the same amount, using the same method. Paid plans let you do it more often, and add the tuned adapter, the synced lesson store and team controls.

This matters for a specific reason. The percentage on the homepage is what the free tier produces. If Free metered the reduction itself, the number you could verify would never be the number we published, and the invitation to check our arithmetic would be worthless.

The percentage still changes from one project to another. A large monorepo and a small project behave differently, which is why the dashboard shows your own figure rather than ours.

Checkout and billing management are hosted by Stripe. The extension refreshes a short-lived signed entitlement about daily and remains offline-tolerant. Ending a subscription never deletes local project data.

What context savings can be worth

In one measured session on this repository, Latorium kept 648,692 input tokens out of the agent’s context. The dollar value depends on the model and how those tokens would otherwise have been billed.

Model List price per million input tokens One session Twenty sessions
Claude Opus 5 $5.00 $3.24 $64.87
Claude Sonnet 5 $2.00 $1.30 $25.95
OpenAI GPT-5, via Codex $2.50 $1.62 $32.43

The table applies list prices to the measured token count. It is an example, not a bill forecast.

List prices checked 19 September 2026. Cached input may cost less than new input, and provider prices change — Anthropic cut Opus from $15.00 to $5.00 per million input tokens, which cut the figures in this table by roughly two thirds. Check current rates before relying on it. The dashboard uses your own project data; this table uses one Latorium development session.

1M-token example

Retrieval reduces input. Local delegation can reduce paid output. The charts separate the two.

Input your agent no longer has to read

Across 204 retrievals on this repository, 77.5% of the text in the files retrieval selected never reached the agent. That is the denominator: the complete text of the files chosen for each retrieval, not the repository, and not files retrieval never opened. The comparison below uses a rounder 75%.

Claude Opus 5$5.00 list$3.75 saved
Claude Sonnet 5$2.00 list$1.50 saved
GPT-5 via Codex$2.50 list$1.88 saved
still sent kept out

Output Local Coder can write instead

Retrieval does not reduce output tokens. Delegation does. Summaries, scaffolding, boilerplate and first drafts can run locally instead. The comparison below assumes 70% of that routine output is delegated — a setting you control, not a result we claim.

Claude Opus 5$25.00 list$17.50 saved
Claude Sonnet 5$10.00 list$7.00 saved
GPT-5 via Codex$20.00 list$14.00 saved
still bought from the provider drafted locally, free

Output costs five times input or more, so delegated work has more leverage. At 70% delegation, an Opus 5 user keeps $17.50 of every $25.00 of routine output.

How these numbers are calculated Input savings are measured on this repository. Output savings are arithmetic from an assumed 70% delegation rate, not a benchmark result. Delegate nothing and output savings are zero. Provider prices also change, so check current rates before relying on either figure.

Install Latorium in VS Code

Install from the Marketplace and let Latorium set up the rest.

Latorium

Published VS Code extension

VS Code 1.95+Windows · macOS · Linux
Install for VS Code
  1. Install from the Marketplace, or open Extensions in VS Code and search for Latorium.
  2. Open Latorium from the VS Code sidebar. If setup has not run, the health line at the top offers Set up local AI. It shows the download size and disk space before starting.
  3. Wait for Ready to code. Setup installs the required local component, downloads the model, and checks that it can answer.
  4. To use Latorium with Codex or Claude Code, open Advanced tools at the bottom of the sidebar, choose Connect Codex & Claude, then restart that assistant once — it reads MCP configuration at startup.

If setup stops, run Latorium: Set Up or Repair from the Command Palette. Existing model downloads are kept, so retrying does not start from zero. No terminal command is required.

Install from the terminal instead
code --install-extension latorium.latorium
4 GB VRAM
~20 tokens/sec
8 GB VRAM
~40 tokens/sec
CPU only
Works, slower
Disk, 4K to 12K context
5.97 GB to 7.20 GB
Embedding model, optional
0.27 GB extra
System memory
16 GB
VS Code
1.95 or later

Every context window — 4K, 6K, 8K and 12K — is available on every plan, including Free. The window is a speed and memory decision, not something we sell you. Measured at the shipped quantisation with a display attached: 12K context runs on a 4 GB card at roughly 20 tokens a second, and 8 GB roughly doubles that. A GPU is optional.

A larger window costs residency. On one card, GPU residency fell from 70% at 4K to 57% at 12K, so the model spills progressively more to system memory as you widen it. If generation feels slow, drop a step before you blame the card — the window is a dropdown in the sidebar and you can change it at any time. Air-gapped install guide