RAG Recall Eval
RAG service that proves its own retrieval : recall@3 = 0.886, MRR@3 = 0.805 with offline stdlib TF-IDF retriever.
- 0.886recall@3over 35 labeled Q's
- 0.805MRR@3
- 1.00Mean faithfulness
- 0 / 35Hallucination flags
AI · MULTI-AGENT · 31-gate11 AGENTS · 1 ORCHESTRATOR · 6 PROJECTS
CHRISTIAN.T.MACIONUTC+811 AGENTS31-gateOWNER-VERIFIED
Every LLM output I ship passes the same statistical gate stack as a systematic strategy. The harness is the artifact; the model is interchangeable.
The platform
4 inputs feed the orchestrator. The orchestrator dispatches to RAG-Recall, Eval-Gate, and Router. The middles publish to Artifact, Metric, and Feedback. Feedback loops back to the orchestrator. that loop is the entire system.
Live graph · public-safe slice
Every claim, every artifact, every agent. wired through one knowledge graph that ships at compile time. T1 (5px) > T2 (3.5px) > T3 (2.5px). Squad color marks the channel. NDA-protected entities are filtered out before public.
BY CHRISTIAN MACION · QUANT RESEARCH + AI ENGINEERING
- agents online · - edges
Loop engineering
Loop engineering is what the AI lane is. The first step is the harness, not the prompt. The last step is the gate, not the model's confidence.
Draft the eval harness BEFORE writing the prompt.
Score against G1 to G31. Evidence per gate, exit codes.
Tier-routed agents. Per-step budget. Cut-over only on pass.
OTel traces · Cohen's κ · bootstrap CIs · DSR.
Reflect → revise → loop. No-progress halt ships nothing.
Beyond the loop · STELLA Office
The 5-rung spine to prompt → skills → agents → loops → graphsn/a is what STELLA's 11-agent office runs on. The graph layer is the productivity unlock: typed entities + relations + provenance, queryable, shared across every agent so the corpus isn't re-derived on every dispatch.
A
11 agents in 6 sub-teams. 4 doctrinal corpora. 5-must-have compliance on every loop primitive. NDA-safe by construction. counts and taxonomy only, never officer names.
Methodology →B
9 entity types. 9 predicates. Four-stage pipeline. Extract · Resolve · Assemble · Query. The AAR → graph contract: every mission writes its findings, every AAR mutates the graph.
Workbook trilogy →C
Most teams plateau at the loop. STELLA runs at the graph rung. and ships the public-facing narrative via two workbooks (AI Engineering from the Ground Up, Graph Engineering for Everyone).
Read W1 →Projects · the surface evidence
Every project on this page ships with a reproducible harness, a labeled eval set or scored output, and a CI-enforced gate stack. None depend on a paid API at read time.
RAG service that proves its own retrieval : recall@3 = 0.886, MRR@3 = 0.805 with offline stdlib TF-IDF retriever.
ReAct-style tool-calling agent with OTel traces, fault injection, and 100% tool/arg correctness.
LLM-as-judge pipeline validated against human raters : Cohen's κ = 0.58 with bootstrap CI and position-bias measured.
MCP server exposing the slop-evaluation gate over Tools, Resources, and Prompts : 20/20 conformance, 100% round-trip parity.
Reflection-loop agent : mean SLOP 127.5 → 14.0 across drafts, 3/4 improved, 1/4 honest no-progress halt.
13-metric literature-grounded AI-output quality gate : drove a real draft from HEAVY (81) to CLEAN (3).
The eval console
Same 31-gate you'd run on a published systematic strategy. Mechanical (exit-0 contract), not opinionated. A pass means the candidate is allowed to surface; a fail means it isn't.
| Gate ID | Label | Result | Value |
|---|---|---|---|
| [G1] | schema conformance | PASS | 100 % |
| [G2] | exit-0 contract | PASS | 0.997 |
| [G3] | structured output parse | PASS | 0.992 |
| [G4] | JSON Schema strict-mode | PASS | 100 % |
| [G5] | byte-length budget | PASS | 2 144 B |
| Gate ID | Label | Result | Value |
|---|---|---|---|
| [G6] | block-bootstrap CIs | CHECK | n=10 000 |
| [G7] | random-timing nulls | PASS | p = 0.41 |
| [G8] | regime-shuffle | PASS | p = 0.62 |
| [G9] | parameter-prior sensitivity | PASS | 0.318 |
| [G10] | Monte-Carlo coverage | PASS | 0.94 |
| Gate ID | Label | Result | Value |
|---|---|---|---|
| [G11] | Locked OOS windows | CHECK | 24 mo |
| [G12] | 5-era stability | PASS | 4 / 5 |
| [G13] | expanding / rolling walk-fwd | PASS | 6 folds |
| [G14] | frozen-spec evaluation | PASS | 3 rev |
| [G15] | embargoed test set | PASS | holdout |
| Gate ID | Label | Result | Value |
|---|---|---|---|
| [G16] | Deflated Sharpe Ratio (DSR) | PASS | 0.70 |
| [G17] | CSCV-based PBO | PASS | 0.18 |
| [G18] | Minimum Backtest Length (MinBTL) | PASS | 8.4 yr |
| [G19] | Bonferroni to Holm | PASS | p = 0.04 |
| [G20] | BH-FDR | PASS | q = 0.07 |
| Gate ID | Label | Result | Value |
|---|---|---|---|
| [G21] | point-in-time dataset | CHECK | refresh Q |
| [G22] | survivorship-bias check | PASS | verified |
| [G23] | rebalance vs signal timing | PASS | t-1 |
| [G24] | frozen-spec (no leak) | PASS | enforced |
| [G25] | no leakage to scoring | PASS | verified |
| Gate ID | Label | Result | Value |
|---|---|---|---|
| [G26] | spread / slippage / latency | FAIL | 7.2 bps |
| [G27] | capacity-constraint report | PASS | $250 k |
| [G28] | funding-carry | CHECK | 1.42 % |
| [G29] | borrow cost | PASS | 0.18 % |
| [G30] | vol-targeted sizing | PASS | σ → 8 % |
| [G31] | regime overlay | PASS | ON |
Open-source on GitHub · 31-gate enforced by CI · same gate stack as a published systematic strategy · stdlib-only Python at the core · the one FAIL on G26 is real. slippage budget overrun caught by the harness, not a model that guessed.
JUMP TO · ⌘K
LIVE COVERAGE · L1 / L2 / L3 · ⌘J
drag to rotate · hover pauses
data · Yahoo 6 + CoinGecko 4 + ECB FX + GDELT 15m · L1 = top-of-book · L2 = consolidated tape · L3 = full depth · 12 venues · 5 continents · UTC+8 home base