PUBLIC DEVELOPMENT ROUTE · UPDATED 15 AUG 2026

What is built and what still needs work

The v0.50.0 foundation now includes guided setup, production accounts, public-repository benchmarking, signed model updates, and a hardened release package. The next work is Cloud Boost, broader hardware evidence, and turning research tooling into supported product features.

Current release
v0.50.0
Verification
685 / 686 tests
Adversarial suite
20 / 20 passed
Context model
Evaluation passed

These figures describe the current development build. They will be updated as new release checks and hardware results are added.

01
SHIPPED · AUG 2026 Shipped

Retrieval foundation

Retrieval foundation is in the product.

The original retrieval phase landed early: large-repository discovery, persistent graph retrieval, iterative retrieval, dynamic context control, confidence, token budgeting, compression, task-specific strategies and inspection controls are implemented.

  1. Remove the 500-file ceiling

    Make workspace discovery and indexing cover large repositories cleanly instead of stopping at the current search limit.

    Shipped
  2. Persistent code graph

    Track imports, calls, references, inheritance, interfaces, implementations and test relationships as first-class repository structure.

    Shipped
  3. Hybrid semantic + graph retrieval

    Use embeddings to find strong candidates, then expand through real code relationships instead of choosing semantic or structural ranking in isolation.

    Shipped
  4. Iterative retrieval

    Let the model ask for a specific missing caller, test, symbol or file halfway through a task instead of front-loading everything.

    Shipped
  5. Dynamic context eviction

    Treat context like working memory: load it, use it, discard it, and replace it as the task moves on.

    Shipped
  6. Retrieval confidence

    Estimate when the current evidence is enough. Stay lean when confidence is high and widen the search when it is not.

    Shipped
  7. Unified token-aware budgeting

    Replace scattered character caps with one allocator that understands the total context budget and the value of each source.

    Shipped
  8. Token-aware reranking

    Prefer a short, highly useful function over a huge file that scores only slightly better.

    Shipped
  9. Duplicate-context removal

    Detect overlapping chunks, repeated definitions, generated noise and information the conversation already contains.

    Shipped
  10. Context compression

    Use signatures, types or compact summaries for lower-priority material, then expand to source only when needed.

    Shipped
  11. Task-specific retrieval plans

    Debugging should look for failures and callers; refactors for references and interfaces; tests for behavior and nearby test patterns.

    Shipped
  12. Context Inspector

    Show exactly what was retrieved, why it ranked, its token cost and what relationship pulled it into the packet.

    Shipped
  13. Pin and exclude controls

    Let developers force important context in, or keep generated and irrelevant files out.

    Shipped
Gate RETRIEVAL v2
02
AHEAD OF OCT–NOV 2026 TARGET In validation

Scale & speed

Architecture shipped; representative hardware and repository evidence is the remaining gate.

The benchmark matrix, routing evidence, hardware profiling, lazy analysis, parallel preprocessing, Git awareness, and telemetry are implemented. More machines and repositories still need testing.

  1. Large-repository benchmark mode

    Stress retrieval and indexing on thousands of files and repositories measured in hundreds of thousands or millions of lines. The matrix runner now automates bounded repository × model × repeated-run execution.

    Validate
  2. Weak-hardware benchmark mode

    Measure CPU-only inference, low VRAM, memory pressure and model spill instead of assuming a gaming GPU. The tooling exists; broader weak-hardware runs still need to be collected.

    Validate
  3. Hardware-aware context sizing

    Choose a sensible context ceiling from observed throughput, memory pressure and model fit.

    Shipped
  4. VRAM/RAM/model-fit detection

    Detect when a model is comfortable, marginal, or falling back to slower execution. RAM/hardware profiling exists; GPU/VRAM detection without a loaded Ollama model still needs strengthening.

    Improve
  5. Automatic model recommendations

    Choose the strongest compatible signed model release from detected RAM, ask before downloading, smoke-test it, and switch only after it works.

    Shipped
  6. Empirical model routing

    Route by historical task success, latency and resource cost, not only a hand-written complexity rule.

    Shipped
  7. Dedicated completion model

    Keep inline prediction on the smallest model that is actually good enough, separate from heavier reasoning work.

    Shipped
  8. Speculative completion cache

    Prepare likely completions while the developer types and invalidate them cheaply when the code changes direction.

    Shipped
  9. Lazy deep analysis

    Build cheap repository metadata first and delay expensive work until a task needs it.

    Shipped
  10. Parallel preprocessing

    Run safe independent work such as diagnostics, metadata and retrieval preparation concurrently.

    Shipped
  11. Git-diff awareness

    Give recent edits and active changes more weight when they are relevant to the task.

    Shipped
  12. Per-stage profiling

    Separate indexing, retrieval, prompt evaluation, generation and verification time so optimization targets the real bottleneck.

    Shipped
  13. Time-to-first-token telemetry

    Measure retrieval overhead separately from model prompt evaluation and generation.

    Shipped
Gate SCALE-READY
03
AHEAD OF DEC 2026–JAN 2027 TARGET Shipped + hardening

Safer agent loops

Agent safety and recovery are implemented and still being hardened.

Impact analysis, risk scoring, adaptive verification, repair-loop stopping, regression tracking, multi-step planning, rollback, project-state awareness, Ollama recovery and sandbox hardening are implemented. v0.50.0 also runs all twenty adversarial catalog scenarios into executable end-to-end checks.

  1. Impact analysis before edits

    Estimate which symbols, files and tests an autonomous change can affect before writing anything.

    Shipped
  2. Change-risk scoring

    Distinguish an isolated edit from a cross-module change and scale verification accordingly.

    Shipped
  3. Adaptive verification

    Run the cheapest meaningful checks first, then broaden only when the risk or failures justify it.

    Shipped
  4. Repair-loop stopping

    Recognize when retries are no longer improving the failure and stop wasting inference on the same dead end.

    Shipped
  5. Stronger regression tracking

    Track whether a repair strategy later creates problems, not just whether it passed once.

    Shipped
  6. Multi-step verified planning

    Break larger work into smaller goals with verification between them instead of one opaque edit pass.

    Shipped
  7. Partial rollback

    Keep verified successful substeps when a later part fails instead of throwing away the whole task.

    Shipped
  8. Project-state awareness

    Plan around branch, dirty files, recent failures and the current test state.

    Shipped
  9. Ollama failure hardening

    Handle missing models, server loss, timeouts, partial streams and model changes without leaving the task in a strange state.

    Shipped
  10. Sandbox hardening

    Make isolation choices clearer and safer for autonomous validation.

    Shipped
  11. Adversarial safety suite

    Continuously test malformed edits, path traversal, protected files, interrupted operations and hostile model output. The full catalog now runs as executable end-to-end checks in CI.

    20 / 20
Gate VERIFIED AGENT
04
175 DAYS BEFORE ORIGINAL START Substantially shipped

Learning & memory

Learning and memory moved from 2027 into the August 2026 build.

Retrieval-policy learning, stale invalidation, confidence/decay, editable and scoped memory, knowledge export/import, team packs, dataset building and quality filtering are already implemented. Remaining work is deeper seed-playbook validation and a stronger versioned dataset pipeline.

  1. Retrieval-policy learning

    Record which context choices led to verified success and favor those strategies on similar future tasks.

    Shipped
  2. Stale-memory invalidation

    Tie lessons to file/chunk hashes or repository state so major code changes lower their confidence automatically.

    Shipped
  3. Memory confidence and decay

    Let old or repeatedly unhelpful lessons lose influence instead of living forever at full strength.

    Shipped
  4. User-editable memory

    Inspect, pin, correct, disable or delete what Latorium has learned.

    Shipped
  5. Project vs global knowledge

    Separate facts about one repository from reusable programming lessons that can safely apply elsewhere.

    Shipped
  6. Knowledge export/import

    Move approved Latorium memory between machines without moving an entire workspace.

    Shipped
  7. Team knowledge packs

    Share reviewed conventions and lessons without turning private project code into a shared memory dump.

    Shipped
  8. Seed repair expansion

    Add more built-in repair playbooks, but only where reproducible tests show they help. The framework exists; the repair knowledge base still needs broader validated coverage.

    Expand
  9. Seed-data evaluation

    Benchmark every shipped lesson so weak generalizations do not become defaults for every install. Validation exists; stronger fixtures and more empirical evidence remain.

    Expand
  10. Training dataset builder

    Turn verified tasks into structured examples that can later be used for local neural personalization.

    Shipped
  11. Training-data quality filters

    Reject ambiguous outcomes, stale examples, accidental secrets and examples with too much irrelevant context.

    Shipped
Gate LEARNING LOOP
05
AFTER DATASET + HARDWARE VALIDATION Research

Local learning R&D

Neural specialization remains optional research, not a launch dependency.

The training scripts, comparison gate, and a small context-reranker prototype now exist. They remain R&D rather than customer features: adapter training still needs a reviewed dataset and supported CUDA hardware, and the reranker has not been integrated into retrieval.

  1. Local LoRA training

    A reproducible QLoRA pipeline is implemented, but no customer adapter is shipped and real training remains evidence- and hardware-gated.

    Pipeline built
  2. LoRA evaluation gate

    A regression gate compares base and adapter results and rejects correctness loss, excessive latency, or insufficient samples.

    Gate built
  3. Automatic adapter rollback

    The gate emits an immutable rejection and rollback decision; automatic installation and rollback of customer adapters are not productized.

    Decision built
  4. Latorium context-selection model

    A small cross-encoder prototype passed its held-out local evaluation; broader datasets and production retrieval integration remain.

    Prototype
  5. Latorium-specialized adapter/model

    Explore a model trained around Latorium's retrieval, verification and tool-use workflow instead of generic chat behavior.

    Research
Gate R&D GATE
06
CONTINUOUS · v0.50.0 Active

Release hardening

Release verification is now a real gate, not a checklist.

The release gate now covers 20 executable adversarial checks, package secret scanning, benchmark retention, minified source-map-free bundles, signed model releases, and recovery tests for the main persistent indexes. Clean-checkout reproducibility and broader native GPU detection remain.

  1. Broader persistent-store recovery tests

    Quarantine and rebuild corrupt or truncated indexes, migrate older schemas, and verify atomic writes and abandoned-lock recovery.

    Shipped
  2. Clean-build verification

    Rebuild from a fresh checkout/environment and require the entire release gate to pass.

    Next
  3. Reproducible VSIX verification

    Compare deterministic package manifests and hashes so release contents are explainable and repeatable.

    Next
  4. GPU/VRAM detection without a loaded model

    Probe hardware independently of Ollama runtime state, then merge native detection with observed model evidence.

    Next
  5. Hardware & routing inspector

    Show detected profiles, imported evidence, model-fit estimates and the reason behind each routing decision.

    Next
  6. Automatic model install/fit recommendations

    Offer guided runtime setup and signed, RAM-aware model releases with explicit download consent and a post-download smoke test.

    Shipped
  7. Benchmark regression release gate

    Keep benchmark regressions attached to release verification and surface them in the dashboard.

    Shipped
  8. Adversarial CI suite

    Run all twenty catalog scenarios cross-platform as part of verification.

    20 / 20
  9. Benchmark matrix runner

    Run bounded repository × model × repeated-run matrices sequentially to avoid VRAM thrashing.

    Shipped
  10. Benchmark history retention & pruning

    Bound history by age, count and storage, with dry-run pruning before deletion.

    Shipped
  11. Experimental routing lane

    Run benchmark-driven routing experiments without silently mutating stable routing policy.

    Next
  12. Dataset versioning and quality reports

    Add versioning, deduplication, balancing, leakage controls and build reports to the training-data pipeline.

    Next
  13. LoRA training experiments

    The reproducible QLoRA trainer and regression-based adapter gate exist, but real training still requires a reviewed dataset, supported CUDA hardware, and repeatable evaluation evidence.

    Research
Gate RELEASE
Latorium