WORK · WORKING ARCHIVECASE STUDIES · SHIPPED RECORD · WHAT SURVIVED REVIEW

CHRISTIAN.T.MACIONUTC+8CASE STUDIESOWNER-VERIFIED

A working archive of what shipped.

Trade-offs, receipts, and the parts that survived review.

The portfolio index for the work-as-shipped. Quant-research cases across five asset classes; AI-engineering cases with reproducible scorecards; the live production strip; and the frontier: graph engineering, pedagogy, and the doctrine corpus that anchors the voice.

01 / the work loop

Hypothesis → backtest → paper → live → postmortem.

The work runs on a five-stage loop. A hypothesis is registered with a pre-declared fire rule and a frozen-spec evaluation window. The backtest is run behind the 31-gate statistical filter to compute block-bootstrap CIs, deflated Sharpe, CSCV-based PBO, walk-forward stability, and the multiple-testing layer that prevents the funnel from manufacturing fake winners. Survivors move to paper trading with a capped notional and a pre-registered risk overlay. A live leg is opened only after the paper leg produces un-gameable forward out-of-sample evidence. The postmortem is always written with KILLs included and banked into the methodology as enforced discipline.

The AI-engineering work follows the same loop with different primitives. A prompt or scaffold is the hypothesis. The eval harness (faithfulness, hallucination flags, retrieval recall@k) is the backtest. A staging deployment with a frozen eval set is the paper leg. Production rollout behind a kill-switch is the live leg. The postmortem captures the failure modes, not the wins, because the wins tend to be self-promoting while the failures compound.

The loop is the same on both sides because the failure mode is the same: an unsupervised optimization that drifts toward whatever produces the prettiest scorecard. The loop is the leash.

02 / quant case studies

Five public-data case studies.

Each card links to a public-data project brief. Every project ships with the data source, the frozen-spec note, the 31-gate gate evidence, and a copy of the methodology it was tested against. None of the projects rely on proprietary data sources. Every line is reproducible on a fresh install.

03 / ai engineering case studies

Three reproducible eval-first cases.

Every AI case ships with a reproducible scorecard. The eval harness is the headline artifact to not the demo. Each project below commits to an in-repo labeled set, runs the harness, and reports the numbers in a single PDF.

  • AI · RAG

    RAG Recall Eval

    RAG service that proves its own retrieval : recall@3 = 0.886, MRR@3 = 0.805 with offline stdlib TF-IDF retriever.

    recall@30.886MRR@30.805
    Read the brief →
  • AI · AGENTS

    Tool-Call Agent

    ReAct-style tool-calling agent with OTel traces, fault injection, and 100% tool/arg correctness.

    Tool / arg correctness100%Injected faults recovered6 / 6
    Read the brief →
  • AI · LLM-AS-JUDGE

    LLM-as-Judge Harness

    LLM-as-judge pipeline validated against human raters : Cohen's κ = 0.58 with bootstrap CI and position-bias measured.

    Cohen's κ (vs human)0.58Pass rate0.60 ± 0.02
    Read the brief →

04 / live in production

What is running right now.

Production surfaces that survive the omit-this-if-stale rule. Each is byte-deterministic at build time; refresh cadence is on the page itself.

LIVE

Markets terminal

Index tape, order-book depth, term structure, coverage globe. BTC-USD, S&P, DXY, VIX. Updated at build.

Open the terminal →

LIVE

Prediction markets

Eight binary event contracts with implied probability, edge, and a 12-point sparkline. NDA-safe by construction.

Open the book →

LIVE

Now

Dated one-pager. Open-to / not-open-to / where / how to reach. Updated ~monthly, on job-change or regime-event override.

Open the page →

05 / what ships next

Three solution tracks compounding right now.

The shipped record is downstream of the build rhythm. Three workstreams compound upstream. Each ships with a real artifact you can verify. not a claim backed only by a name. The discipline is the same in every lane: ship behind a gate, widen the funnel second, log the postmortem.

Quant lane. The statistical-arb pipeline grows toward live deployment with new asset classes (FX, commodities), tighter gate thresholds, and a regime-aware deployment layer. The compounding asset is the postmortem library. every KILL becomes capital for the next candidate.

AI lane. The eval harness grows the same way. Every shipped agent adds a frozen eval set, a kill-switch, and a behavioral test plan. The compounding asset is the eval-set library. every failure pattern is now a regression test for the next build.

Build-discipline lane. The pattern that survives a regime flip is the loop itself: hypothesis → backtest → paper → live → postmortem. Add new disciplines rather than stuffing existing ones. The next twelve months sit roughly across these three lanes, with the work loop as the constant.

For the eval-first methodology → see the 31-gate filter.

The methodology page documents the eval gate.

Last build: 2026-08-09 · Quant Researcher · 31-gate eval-first