PUBLIC DEVELOPMENT ROUTE · UPDATED 15 AUG 2026
What is built and what still needs work
The v0.50.0 foundation now includes guided setup, production accounts, public-repository benchmarking, signed model updates, and a hardened release package. The next work is Cloud Boost, broader hardware evidence, and turning research tooling into supported product features.
- Current release
- v0.50.0
- Verification
- 685 / 686 tests
- Adversarial suite
- 20 / 20 passed
- Context model
- Evaluation passed
These figures describe the current development build. They will be updated as new release checks and hardware results are added.
Retrieval foundation
Retrieval foundation is in the product.
The original retrieval phase landed early: large-repository discovery, persistent graph retrieval, iterative retrieval, dynamic context control, confidence, token budgeting, compression, task-specific strategies and inspection controls are implemented.
-
Remove the 500-file ceilingShipped
Make workspace discovery and indexing cover large repositories cleanly instead of stopping at the current search limit.
-
Persistent code graphShipped
Track imports, calls, references, inheritance, interfaces, implementations and test relationships as first-class repository structure.
-
Hybrid semantic + graph retrievalShipped
Use embeddings to find strong candidates, then expand through real code relationships instead of choosing semantic or structural ranking in isolation.
-
Iterative retrievalShipped
Let the model ask for a specific missing caller, test, symbol or file halfway through a task instead of front-loading everything.
-
Dynamic context evictionShipped
Treat context like working memory: load it, use it, discard it, and replace it as the task moves on.
-
Retrieval confidenceShipped
Estimate when the current evidence is enough. Stay lean when confidence is high and widen the search when it is not.
-
Unified token-aware budgetingShipped
Replace scattered character caps with one allocator that understands the total context budget and the value of each source.
-
Token-aware rerankingShipped
Prefer a short, highly useful function over a huge file that scores only slightly better.
-
Duplicate-context removalShipped
Detect overlapping chunks, repeated definitions, generated noise and information the conversation already contains.
-
Context compressionShipped
Use signatures, types or compact summaries for lower-priority material, then expand to source only when needed.
-
Task-specific retrieval plansShipped
Debugging should look for failures and callers; refactors for references and interfaces; tests for behavior and nearby test patterns.
-
Context InspectorShipped
Show exactly what was retrieved, why it ranked, its token cost and what relationship pulled it into the packet.
-
Pin and exclude controlsShipped
Let developers force important context in, or keep generated and irrelevant files out.
Scale & speed
Architecture shipped; representative hardware and repository evidence is the remaining gate.
The benchmark matrix, routing evidence, hardware profiling, lazy analysis, parallel preprocessing, Git awareness, and telemetry are implemented. More machines and repositories still need testing.
-
Large-repository benchmark modeValidate
Stress retrieval and indexing on thousands of files and repositories measured in hundreds of thousands or millions of lines. The matrix runner now automates bounded repository × model × repeated-run execution.
-
Weak-hardware benchmark modeValidate
Measure CPU-only inference, low VRAM, memory pressure and model spill instead of assuming a gaming GPU. The tooling exists; broader weak-hardware runs still need to be collected.
-
Hardware-aware context sizingShipped
Choose a sensible context ceiling from observed throughput, memory pressure and model fit.
-
VRAM/RAM/model-fit detectionImprove
Detect when a model is comfortable, marginal, or falling back to slower execution. RAM/hardware profiling exists; GPU/VRAM detection without a loaded Ollama model still needs strengthening.
-
Automatic model recommendationsShipped
Choose the strongest compatible signed model release from detected RAM, ask before downloading, smoke-test it, and switch only after it works.
-
Empirical model routingShipped
Route by historical task success, latency and resource cost, not only a hand-written complexity rule.
-
Dedicated completion modelShipped
Keep inline prediction on the smallest model that is actually good enough, separate from heavier reasoning work.
-
Speculative completion cacheShipped
Prepare likely completions while the developer types and invalidate them cheaply when the code changes direction.
-
Lazy deep analysisShipped
Build cheap repository metadata first and delay expensive work until a task needs it.
-
Parallel preprocessingShipped
Run safe independent work such as diagnostics, metadata and retrieval preparation concurrently.
-
Git-diff awarenessShipped
Give recent edits and active changes more weight when they are relevant to the task.
-
Per-stage profilingShipped
Separate indexing, retrieval, prompt evaluation, generation and verification time so optimization targets the real bottleneck.
-
Time-to-first-token telemetryShipped
Measure retrieval overhead separately from model prompt evaluation and generation.
Safer agent loops
Agent safety and recovery are implemented and still being hardened.
Impact analysis, risk scoring, adaptive verification, repair-loop stopping, regression tracking, multi-step planning, rollback, project-state awareness, Ollama recovery and sandbox hardening are implemented. v0.50.0 also runs all twenty adversarial catalog scenarios into executable end-to-end checks.
-
Impact analysis before editsShipped
Estimate which symbols, files and tests an autonomous change can affect before writing anything.
-
Change-risk scoringShipped
Distinguish an isolated edit from a cross-module change and scale verification accordingly.
-
Adaptive verificationShipped
Run the cheapest meaningful checks first, then broaden only when the risk or failures justify it.
-
Repair-loop stoppingShipped
Recognize when retries are no longer improving the failure and stop wasting inference on the same dead end.
-
Stronger regression trackingShipped
Track whether a repair strategy later creates problems, not just whether it passed once.
-
Multi-step verified planningShipped
Break larger work into smaller goals with verification between them instead of one opaque edit pass.
-
Partial rollbackShipped
Keep verified successful substeps when a later part fails instead of throwing away the whole task.
-
Project-state awarenessShipped
Plan around branch, dirty files, recent failures and the current test state.
-
Ollama failure hardeningShipped
Handle missing models, server loss, timeouts, partial streams and model changes without leaving the task in a strange state.
-
Sandbox hardeningShipped
Make isolation choices clearer and safer for autonomous validation.
-
Adversarial safety suite20 / 20
Continuously test malformed edits, path traversal, protected files, interrupted operations and hostile model output. The full catalog now runs as executable end-to-end checks in CI.
Learning & memory
Learning and memory moved from 2027 into the August 2026 build.
Retrieval-policy learning, stale invalidation, confidence/decay, editable and scoped memory, knowledge export/import, team packs, dataset building and quality filtering are already implemented. Remaining work is deeper seed-playbook validation and a stronger versioned dataset pipeline.
-
Retrieval-policy learningShipped
Record which context choices led to verified success and favor those strategies on similar future tasks.
-
Stale-memory invalidationShipped
Tie lessons to file/chunk hashes or repository state so major code changes lower their confidence automatically.
-
Memory confidence and decayShipped
Let old or repeatedly unhelpful lessons lose influence instead of living forever at full strength.
-
User-editable memoryShipped
Inspect, pin, correct, disable or delete what Latorium has learned.
-
Project vs global knowledgeShipped
Separate facts about one repository from reusable programming lessons that can safely apply elsewhere.
-
Knowledge export/importShipped
Move approved Latorium memory between machines without moving an entire workspace.
-
Team knowledge packsShipped
Share reviewed conventions and lessons without turning private project code into a shared memory dump.
-
Seed repair expansionExpand
Add more built-in repair playbooks, but only where reproducible tests show they help. The framework exists; the repair knowledge base still needs broader validated coverage.
-
Seed-data evaluationExpand
Benchmark every shipped lesson so weak generalizations do not become defaults for every install. Validation exists; stronger fixtures and more empirical evidence remain.
-
Training dataset builderShipped
Turn verified tasks into structured examples that can later be used for local neural personalization.
-
Training-data quality filtersShipped
Reject ambiguous outcomes, stale examples, accidental secrets and examples with too much irrelevant context.
Local learning R&D
Neural specialization remains optional research, not a launch dependency.
The training scripts, comparison gate, and a small context-reranker prototype now exist. They remain R&D rather than customer features: adapter training still needs a reviewed dataset and supported CUDA hardware, and the reranker has not been integrated into retrieval.
-
Local LoRA trainingPipeline built
A reproducible QLoRA pipeline is implemented, but no customer adapter is shipped and real training remains evidence- and hardware-gated.
-
LoRA evaluation gateGate built
A regression gate compares base and adapter results and rejects correctness loss, excessive latency, or insufficient samples.
-
Automatic adapter rollbackDecision built
The gate emits an immutable rejection and rollback decision; automatic installation and rollback of customer adapters are not productized.
-
Latorium context-selection modelPrototype
A small cross-encoder prototype passed its held-out local evaluation; broader datasets and production retrieval integration remain.
-
Latorium-specialized adapter/modelResearch
Explore a model trained around Latorium's retrieval, verification and tool-use workflow instead of generic chat behavior.
Release hardening
Release verification is now a real gate, not a checklist.
The release gate now covers 20 executable adversarial checks, package secret scanning, benchmark retention, minified source-map-free bundles, signed model releases, and recovery tests for the main persistent indexes. Clean-checkout reproducibility and broader native GPU detection remain.
-
Broader persistent-store recovery testsShipped
Quarantine and rebuild corrupt or truncated indexes, migrate older schemas, and verify atomic writes and abandoned-lock recovery.
-
Clean-build verificationNext
Rebuild from a fresh checkout/environment and require the entire release gate to pass.
-
Reproducible VSIX verificationNext
Compare deterministic package manifests and hashes so release contents are explainable and repeatable.
-
GPU/VRAM detection without a loaded modelNext
Probe hardware independently of Ollama runtime state, then merge native detection with observed model evidence.
-
Hardware & routing inspectorNext
Show detected profiles, imported evidence, model-fit estimates and the reason behind each routing decision.
-
Automatic model install/fit recommendationsShipped
Offer guided runtime setup and signed, RAM-aware model releases with explicit download consent and a post-download smoke test.
-
Benchmark regression release gateShipped
Keep benchmark regressions attached to release verification and surface them in the dashboard.
-
Adversarial CI suite20 / 20
Run all twenty catalog scenarios cross-platform as part of verification.
-
Benchmark matrix runnerShipped
Run bounded repository × model × repeated-run matrices sequentially to avoid VRAM thrashing.
-
Benchmark history retention & pruningShipped
Bound history by age, count and storage, with dry-run pruning before deletion.
-
Experimental routing laneNext
Run benchmark-driven routing experiments without silently mutating stable routing policy.
-
Dataset versioning and quality reportsNext
Add versioning, deduplication, balancing, leakage controls and build reports to the training-data pipeline.
-
LoRA training experimentsResearch
The reproducible QLoRA trainer and regression-based adapter gate exist, but real training still requires a reviewed dataset, supported CUDA hardware, and repeatable evaluation evidence.
Latorium