AI & Research Framework · Lesson 02 of 05
Building your data stack
What it is
Your data stack is the set of sources and tools that feed step two of the research loop — evidence. It's your workbench: where you actually go to check funding, pull a TVL chart, read the ETF flows, or confirm what real yields are doing.
Here's the reframe that makes this lesson simple: you already know what every tool measures. Thirty-six lessons taught you the signals. The stack is just where each one lives. So instead of a pile of apps, think of six labeled drawers — one per discipline you've studied:
- Structure & price: a charting platform (TradingView is the standard) — ranges, levels, volume profile.
- Derivatives: aggregators like Coinglass, Coinalyze, or Velo — funding, OI, liquidations, basis; Laevitas or Deribit's own data for options skew.
- Flows: Farside or SoSoValue for daily ETF flows; CryptoQuant for exchange flows and the Coinbase Premium.
- On-chain: DefiLlama for TVL and fees (free, canonical); Glassnode or CryptoQuant for supply and holder metrics; a TokenUnlocks-style calendar; Dune for custom questions; a block explorer for ground truth.
- Macro: FRED (the St. Louis Fed's free library) for yields, real rates, money supply; your charting app for DXY and gold; CME FedWatch for the expected rate path.
- The calendars: one economic calendar (CPI, FOMC) and one unlock calendar — the two schedules of known volatility.
Names will change over the years; the drawers won't. Learn the drawer, and swapping tools is trivial.
Why it matters
Beginners fail at evidence in two opposite ways. The first runs on zero data — vibes, influencers, and screenshots — and never knows it, because opinions feel like information. The second drowns in forty dashboards, mistaking tool-collecting for research. Both are avoiding the same hard thing: asking a specific question and checking it.
A good stack has three properties. It's small — one or two tools per drawer, opened when a question requires them, not browsed recreationally. It's verification-first — you know each source's failure modes, because this curriculum taught them: wallet tags are estimates (On-Chain 01), address counts are gameable (02), CVD feeds differ by venue (Derivatives 05), dollar-TVL lies (04). A tool you can't distrust properly is a tool you can't trust properly. And it's reproducible — mostly free, so anyone (including future-you) can re-pull the number and get the same answer. That's not just thrift; on this site it's philosophy: evidence a reader can't check isn't evidence.
One honest note on paid data: it buys convenience, granularity, and history — real value for professionals. It does not buy secret truth. A beginner's edge is judgment, not feeds, and every example in this curriculum was readable on free tiers.
The two readings, always taught together
Read bullish when
- For a method lesson, "bullish" means signs your stack is working: Questions open tools, not the reverse. Your session starts with "is this rally spot-led?" and ends three clicks later — not with an hour of dashboard tourism looking for a feeling.
- Important numbers get cross-checked. Two sources, or one source plus the raw chain, before a number enters your thesis. When providers disagree, you notice — and that disagreement is itself information about the metric's softness.
- Costly-to-fake data anchors your views. Fees over address counts, flows over narratives, positioning over sentiment — the on-chain shelf's filter running as a habit.
- The calendars are checked first. You're never surprised by a CPI print or a token cliff, because known volatility is scheduled, and your week starts by reading the schedule.
Never alone — confirm with the verification triangle & a second source
Read bearish when
- Signs your stack is hurting you: Dashboard hoarding. Twelve tabs, forty indicators, no question. Tool collecting is procrastination in a lab coat — activity that feels like research and produces none.
- Screenshot trust. A cropped, unlabeled chart from social media is not data — no source, no axis, no date, no way to reproduce it. If you can't pull the number yourself, it isn't evidence; it's a rumor with gridlines.
- Single-source certainty. Building a thesis on one provider's estimate of a soft metric (tagged wallets, aggregated CVD) without knowing it's soft. The failure isn't using the source — it's not knowing its error bars.
- Indicator soup. Layering five derived oscillators on one chart until something agrees with you. That's evidence shopping (Lesson 01) with extra steps. The curriculum gave you a few dozen independent lenses — independence, not quantity, is what makes confluence mean anything.
Never alone — confirm with the verification triangle & a second source
Visual explanation
Real market example
August 5, 2024 — one terrifying morning, six drawers, one classified event. Bitcoin crashed from the high-$50,000s to about $49,000 within hours; altcoins fared worse; headlines screamed. A vibes-based observer had a panic. A stack-based one had a procedure.
Macro drawer: the trigger was global — the Bank of Japan had hiked, the yen-carry trade was violently unwinding, and equities worldwide were gapping down. Not a crypto story; a liquidity shock (Macro 01's crunch, in miniature). Derivatives drawer: billions in long liquidations, open interest wiped — a forced-flow cascade with the classic anatomy from Derivatives 04. Flows drawer: spot selling far calmer than the perp carnage; ETF outflows present but modest against the move. On-chain drawer: nothing broken — fees, activity, supply behavior all boring. Structure drawer: the crash swept deep into a higher-timeframe support region and snapped back the same day.
Assembled classification: a macro-triggered leverage flush with intact fundamentals — the type of event that historically recovers, and this one did, within weeks. That's what a stack is for: not predicting the morning, but classifying it correctly by lunchtime while everyone else is still screaming. The drawers turned the scariest candle of the year into a sentence.
How RIX Intel uses this signal
The desk's stack maps to the same six drawers — and two rules govern it. Every published number links its source. That's the evidence-link requirement you see in every Journal publication: reproducibility isn't a courtesy, it's the trust model. And cross-check before publish — soft metrics (tags, aggregations, estimates) get a second source or a raw-chain confirmation before they carry weight in a thesis.
The desk is deliberately tool-agnostic: providers are replaceable, the questions aren't, and free-first sourcing is preferred precisely so readers can verify without a subscription. The stack serves the loop — never the other way around.
Common mistakes
Where this signal ruins people
Collecting tools instead of answering questions. The stack is a means. If you can't name the question a tool answered for you this month, close the tab — it's decoration.
Trusting screenshots. Unverified chart images are the largest single source of false beliefs in crypto. Reproduce it or discard it — there is no third option for evidence.
Single-sourcing soft numbers. Know which metrics are estimates, and never let one provider's guess anchor a thesis alone. Two legs of the triangle, minimum.
Buying data before exhausting free. Paid feeds are convenience, not alpha. If your free-tier process isn't producing insight, the bottleneck is the process — and no subscription fixes that.