WORK · WORKING ARCHIVECASE STUDIES · SHIPPED RECORD · WHAT SURVIVED REVIEW
CHRISTIAN.T.MACIONUTC+8CASE STUDIESOWNER-VERIFIED
A working archive of what shipped.
Trade-offs, receipts, and the parts that survived review.
The portfolio index for the work-as-shipped. Quant-research cases across five asset classes; AI-engineering cases with reproducible scorecards; the live production strip; and the frontier: graph engineering, pedagogy, and the doctrine corpus that anchors the voice.
01 / the work loop
Hypothesis → backtest → paper → live → postmortem.
The work runs on a five-stage loop. A hypothesis is registered with a pre-declared fire rule and a frozen-spec evaluation window. The backtest is run behind the 31-gate statistical filter to compute block-bootstrap CIs, deflated Sharpe, CSCV-based PBO, walk-forward stability, and the multiple-testing layer that prevents the funnel from manufacturing fake winners. Survivors move to paper trading with a capped notional and a pre-registered risk overlay. A live leg is opened only after the paper leg produces un-gameable forward out-of-sample evidence. The postmortem is always written with KILLs included and banked into the methodology as enforced discipline.
The AI-engineering work follows the same loop with different primitives. A prompt or scaffold is the hypothesis. The eval harness (faithfulness, hallucination flags, retrieval recall@k) is the backtest. A staging deployment with a frozen eval set is the paper leg. Production rollout behind a kill-switch is the live leg. The postmortem captures the failure modes, not the wins, because the wins tend to be self-promoting while the failures compound.
The loop is the same on both sides because the failure mode is the same: an unsupervised optimization that drifts toward whatever produces the prettiest scorecard. The loop is the leash.
02 / quant case studies
Five public-data case studies.
Each card links to a public-data project brief. Every project ships with the data source, the frozen-spec note, the 31-gate gate evidence, and a copy of the methodology it was tested against. None of the projects rely on proprietary data sources. Every line is reproducible on a fresh install.
QUANT · MULTIPLE-TESTING
Multiple Testing & the Deflated Sharpe Ratio
Best-of-160 BTC rule: IS Sharpe 1.14 is only 1.24× the pure-noise expectation : DSR = 0.70 (fail).
Read the brief →QUANT · CROSS-SECTIONAL
Cross-Sectional Momentum (18 coins)
IS Sharpe 0.91 → OOS −0.03 : an honest decay; bootstrap CI straddles zero.
Read the brief →QUANT · TIME-SERIES-MOMENTUM
Time-Series Momentum + Vol Targeting
Vol-targeting lifts Sharpe 0.27 → 0.39 and halves max DD (−62% → −30%).
Read the brief →QUANT · VOLATILITY
The Variance Risk Premium (VIX vs Realized)
Implied > realized 85% of 36 yrs; predicts returns, Newey-West t = +6.5.
Read the brief →QUANT · COINTEGRATION
Pairs Trading via Cointegration (BTC / ETH)
ADF −2.11 (not cointegrated), half-life 208 d : the fade loses, as the test predicts.
Read the brief →
03 / ai engineering case studies
Three reproducible eval-first cases.
Every AI case ships with a reproducible scorecard. The eval harness is the headline artifact to not the demo. Each project below commits to an in-repo labeled set, runs the harness, and reports the numbers in a single PDF.
AI · RAG
RAG Recall Eval
RAG service that proves its own retrieval : recall@3 = 0.886, MRR@3 = 0.805 with offline stdlib TF-IDF retriever.
Read the brief →AI · AGENTS
Tool-Call Agent
ReAct-style tool-calling agent with OTel traces, fault injection, and 100% tool/arg correctness.
Read the brief →AI · LLM-AS-JUDGE
LLM-as-Judge Harness
LLM-as-judge pipeline validated against human raters : Cohen's κ = 0.58 with bootstrap CI and position-bias measured.
Read the brief →
04 / live in production
What is running right now.
Production surfaces that survive the omit-this-if-stale rule. Each is byte-deterministic at build time; refresh cadence is on the page itself.
LIVE
Markets terminal
Index tape, order-book depth, term structure, coverage globe. BTC-USD, S&P, DXY, VIX. Updated at build.
Open the terminal →LIVE
Prediction markets
Eight binary event contracts with implied probability, edge, and a 12-point sparkline. NDA-safe by construction.
Open the book →LIVE
Now
Dated one-pager. Open-to / not-open-to / where / how to reach. Updated ~monthly, on job-change or regime-event override.
Open the page →05 / what ships next
Three solution tracks compounding right now.
The shipped record is downstream of the build rhythm. Three workstreams compound upstream. Each ships with a real artifact you can verify. not a claim backed only by a name. The discipline is the same in every lane: ship behind a gate, widen the funnel second, log the postmortem.
Quant lane. The statistical-arb pipeline grows toward live deployment with new asset classes (FX, commodities), tighter gate thresholds, and a regime-aware deployment layer. The compounding asset is the postmortem library. every KILL becomes capital for the next candidate.
AI lane. The eval harness grows the same way. Every shipped agent adds a frozen eval set, a kill-switch, and a behavioral test plan. The compounding asset is the eval-set library. every failure pattern is now a regression test for the next build.
Build-discipline lane. The pattern that survives a regime flip is the loop itself: hypothesis → backtest → paper → live → postmortem. Add new disciplines rather than stuffing existing ones. The next twelve months sit roughly across these three lanes, with the work loop as the constant.
For the eval-first methodology → see the 31-gate filter.
The methodology page documents the eval gate.
Last build: 2026-08-09 · Quant Researcher · 31-gate eval-first