ML & Data Science
Policy BriefRegulation in Force · Policy BriefThe EU AI Act Is Already Enforcing. The Deadline Everyone Watched Was the One That Moved.
The enforcement machinery is running: chatbot disclosure, synthetic-content marking, GPAI documentation demands, and a US state banning algorithmic grocery pricing.
Ibrahim Denis Fofanah·Oct 8, 2026·12 min
Deep DiveML & Data Science · Deep DiveGradient Boosting Still Rules Tabular Data. The Foundation Models Are Gaining Ground.
Four years after a 45-dataset benchmark crowned gradient-boosted trees on tabular data, tabular foundation models (TabPFN, LimiX-2, NVIDIA's Kumo Tabular) are topping leaderboards. What the 2022 verdict got right, where it no longer holds, and how to choose your next tabular model.
Ibrahim Denis Fofanah·Oct 5, 2026·13 min
Benchmark WatchBenchmark Watch · ML & Data Sciencepandas vs Polars vs DuckDB on 3.5 Million Taxi Rides. DuckDB Won Every Query, and the Timings Were the Least Interesting Part.
I ran pandas, Polars, and DuckDB through the same 3.5M-row NYC taxi CSV on a 7GB box. DuckDB won every query and used half the memory of pandas, but the most useful finding was a CSV parsing gotcha, not a timing.
Ibrahim Denis Fofanah·Sep 29, 2026·6 min
TutorialData Engineering · ReliabilityYour Retry Worked. That Is How You Got Two Rows.
Retries fix transient failures, but they can also duplicate payments, records, and side effects. Here is how idempotency keys, unique constraints, and atomic writes make retries safe.
Ibrahim Denis Fofanah·Sep 25, 2026·8 min
- ExplainerAI Foundations · Explained
What Are Embeddings? How Meaning Becomes Math
An embedding turns meaning into a location so that similarity becomes distance. That one idea is the engine under semantic search, recommendations, and the retrieval half of every RAG system, and it is cheap enough to run on your own data.
Ibrahim Denis Fofanah·Sep 25, 2026·7 min
- Deep DiveData Storytelling · Craft
Data Storytelling Teaches You to Present a Finding. Nobody Teaches You to Find One.
Every guide teaches you to present a finding. Almost none teach you to find one. The four questions that break a story open, and a real Sierra Leone trial where every write-up missed it.
Ibrahim Denis Fofanah·Sep 25, 2026·7 min
AnalysisData Engineering · PerformanceOpenAI Got 6× Better CPU Efficiency With Rust. The Lesson Is Why They Waited.
OpenAI’s storage service once pushed more than 20 million requests per second through Python. A Rust rewrite later delivered 6× better CPU efficiency and 15× better memory efficiency. The useful engineering lesson is not “rewrite everything in Rust.” It is why OpenAI knowingly carried the slower system first.
Ibrahim Denis Fofanah·Sep 25, 2026·16 min
- ExplainerFrom Notebook to Production · Part 1
Your Model Works in a Notebook. That Is the Easy Part.
A notebook proves a model can work once. Production requires it to work continuously, on unseen data, within a budget, and to fail safely. The five questions that define the gap, and why deployment is the scarce skill.
Ibrahim Denis Fofanah·Sep 25, 2026·7 min
TutorialSQL · Data QualityYour SQL Query Did Not Fail. NULL Changed the Question.
NULL does not mean zero, blank, or false. It introduces UNKNOWN, and that third truth value can quietly change filters, exclusions, joins, and aggregates.
Ibrahim Denis Fofanah·Sep 25, 2026·10 min
TutorialAnalytics Engineering · Metrics · Semantic LayerYour KPI Needs a Contract: Build Metrics That Survive Dashboards and AI
A metric is more than an aggregation. Define its entity, grain, time, eligibility, joins, tests, and change policy so dashboards and AI return the same answer.
Ibrahim Denis Fofanah·Sep 25, 2026·9 min
TutorialData Engineering · ReliabilityYour Pipeline Passed. Your Dashboard Is Still Wrong.
A green pipeline only proves the code ran. This practical guide shows how contracts, freshness checks, semantic tests, lineage, and safe promotion catch schema drift before it corrupts a dashboard.
Ibrahim Denis Fofanah·Sep 11, 2026·7 min
- TutorialModel Evaluation · Time Series
Your Train-Test Split Is Leaking the Future
A random split can make a forecasting model look excellent by training on records that occur after its test rows. Here is how to build a time-aware evaluation that rehearses production instead of leaking the future.
Ibrahim Denis Fofanah·Sep 7, 2026·10 min