Research
Engineering insights, conference talks, and benchmark results.
Blog Beyond Prompt and Pray: The GMS Methodology, End to End
A methodology walkthrough for model risk and engineering teams — why prompt and pray isn't reliable, the three pillars that replace it and a governed banking-complaint agent built one layer at a time.
Blog Knowlytix 1.0: The Geometric Memory System Is Now Stable
The whole family is live on PyPI at 1.0 — a stable release with a settled API. One command installs it all.
Blog Forge and Proof: Putting AI Agents Into Production You Can Defend
A half-day, hands-on workshop for model risk and compliance leaders. ForgeLoop builds governed agents; ProofLoop proves them — both on the Geometric Memory System.
Paper KAL: Connecting Knowledge Graphs to Geometric Memory Systems
Paper GPU-Accelerated Gradient-Based Space-Filling Design with an Application to Systematic Evaluation of Large Language Models on Financial-Regulatory Documents
Blog How to Test an AI Agent Without Writing a Single Test
If the document already has the answers, why is a human grading the agent? Essay 2 in the series on validating agentic AI systems.
Paper When the Judge is Wrong: Measuring LLM-as-Judge Reliability Against Graph-Verified Ground Truth in Financial Documents
Paper FinStructBench: A Benchmark for Structured Information Retrieval from Financial Documents Using Graph-Verifiable Questions
Blog The Missing Layer Between RAG and Production
RAG solved retrieval. Guardrails solved safety. What's missing is a governed harness — and regulated industries can't ship without it.