Agent Memory Needs Typed Claims, Routing, and Maintenance
A three-month GBrain + Hermes run: why raw transcripts are not memory, and the typed-substrate, routing, and maintenance loop that made durable recall inspectable.
Decision guides for the software and AI you're considering — what works, what it costs, and where it fails.
8 articles
A three-month GBrain + Hermes run: why raw transcripts are not memory, and the typed-substrate, routing, and maintenance loop that made durable recall inspectable.
Letting an agent touch production needs blast-radius limits, not trust. The guardrails to set first, the failure modes they stop, and where to keep a human.
For an internal smart-contract audit SFT lane, hold the run until the dataset passes dedup, leakage, grounding, negative-coverage, and truncation gates.
A shared-host Jentic One install hit release drift, a false health check, and an admin gate before any governed API call; here is what each one costs.
Measured Qwen3.8-27B serving on one 96GB Blackwell GPU: a BF16 vLLM baseline, an NVFP4 SGLang lane, and the throughput and latency that decide which holds.
Receipt-backed LTX 2.3 and ComfyUI workflows show the GPU, retry, and failure-mode trade-offs behind self-hosting versus an occasional hosted API call.
Multi-token prediction cut our latency, then quietly corrupted tool calls. What MTP does, how the failure presented, and the settings that made it safe again.
The Lab tests AI and software buying decisions with a named scenario, a verdict, visible workflow and failure modes, a keep-or-drop rule, and receipts.