Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

3,637 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Diogo Ribeiro

Lead Data Scientist · AI Engineer · Professor · Mathematical Engineer
Working between the United Kingdom and Portugal · Python (typed, NumPy-first) · production AI + reproducible research

Home (current page) Projects Methods Research Teaching

"Knowledge is knowing a tomato is a fruit; wisdom is not putting it in a fruit salad." — Miles Kington

I build production systems that turn complex data into reliable decisions, and reproducible research pipelines that turn open data into auditable evidence. My work spans end-to-end AI systems (LLMs, RAG, and agents), forecasting and anomaly detection, and a growing body of econometric and dynamical-systems research on inequality, wealth, and public policy. Across all of it the constants are the same: lean models, robust software practice, and results you can reproduce and defend.

Poster with the phrase 'Data has a better idea'


What I Work On

  • Production AI & LLM systems
    RAG pipelines, agentic workflows, structured outputs, evaluation loops, and audit-friendly narrative reporting — designed for reliability, observability, and CI from the start.
  • Data science & statistical modelling
    The full toolkit: supervised and unsupervised learning, causal inference and experimentation, survival analysis, Bayesian modelling, and robust/heavy-tailed statistics.
    With a working emphasis on the parts most people skip — uncertainty quantification and calibration, class imbalance, interpretability and fairness, leakage and drift checks, and honest model selection under real-world noise.
  • Forecasting & anomaly detection
    Classical and foundation-model time series (SARIMAX/Prophet through Chronos and masked-patch transformers), conformal prediction intervals, change-point and rare-event detection, and drift monitoring for operational and sensor-driven systems.
  • Econometrics & reproducible policy research
    Panel and causal models, event studies, synthetic control, and Monte Carlo simulation over open economic data (Eurostat, OECD, AMECO, WID, INE/PORDATA, FRED).
  • Dynamical systems & mathematical modelling
    Attractor dynamics, dynamical-systems econometrics, extreme-value and rare-event methods, and exact algorithmic solvers.
  • Data & ML engineering
    Contract-linked ingestion, dataset curation, streaming and lakehouse patterns, and reproducible project scaffolding.

→ Full stack and model families on the Methods tab.


Current Focus

  • Production RAG and agentic systems with evaluation, observability, and CI baked in
  • Reproducible econometric research on inequality, wealth concentration, and public policy — including dynamical-systems models of rentier equilibria and crisis dynamics
  • Robust forecasting, anomaly detection, and drift control for operational data
  • Reusable research infrastructure: contract-linked ETL, GitHub Actions tooling, and reproducibility contracts across multi-paper programmes

Flagship Work

Six projects that cover the range. The full catalogue — around 40 repositories across AI, ML engineering, data engineering, statistics, economics, algorithms, and tooling — is on the Projects tab.

  • feedback-intelligence-agent — Production-style RAG system: a customer feedback intelligence agent with FastAPI, evaluation, observability, and CI.
  • ragops-lab — Evaluation-first RAG and LLMOps platform for production-grade document QA: tracing, regression testing, and cost-aware experimentation.
  • clinic-forecasting-platform — Healthcare demand-forecasting and staffing platform: a 13-model benchmark (SARIMAX, Prophet, gradient boosting, Nixtla, Chronos) with conformal intervals, rolling-origin backtesting, and FastAPI serving.
  • transaction-risk-lakehouse — Production-style PySpark lakehouse for transaction-risk modelling, fraud detection, and temporal model validation.
  • poverty_neoliberalism_research_program — Agent-first scaffold for a ten-paper empirical programme on poverty, wages, taxes, and asset power in the US and UK since 1950, sharing one pipeline and reproducibility contract across all papers.
  • bmssp ⭐ — Deterministic Single-Source Shortest Paths solver for directed graphs with non-negative weights, using a BMSSP-style divide-and-conquer design (typed, tested).

Live dashboards: Portugal Economic Indicators · NASDAQ Stock Analytics


Research, Collaboration & Teaching

I research inequality and political economy, applied econometrics, dynamical systems, production AI, time series and anomaly detection, and survival analysis — mostly organised as multi-paper programmes with a shared pipeline and reproducibility contract. I teach mathematics and data subjects at ESMAD (Instituto Politécnico do Porto) and run seminars on MLOps, streaming analytics, experimentation, and forecasting.

Research & Collaboration for research themes and what I am open to working on.
Teaching for courses, workshops, and course repositories.

When reaching out, include a short note on your use case, constraints, and timeline so we can assess fit quickly.


GitHub Stats

2026 Highlights (public + private work)

  • Production AI: shipped RAG and agent systems (feedback-intelligence-agent, ragops-lab, ai-incident-analysis-agent) — with evaluation, tracing, observability, and CI treated as first-class.
  • ML & data engineering: built serving, MLOps, and lakehouse/streaming platforms (fastapi-ml-platform, feature-store-lab, transaction-risk-lakehouse, pyflink-fraud-detection-streaming, llm-data-platform) spanning inference, drift monitoring, feature parity, and contract-linked ingestion.
  • Research programmes at scale: launched a cluster of reproducible econometric and dynamical-systems projects on inequality, wealth, and policy, several sharing a single pipeline and reproducibility contract across many papers.
  • Algorithms & methods: released and refined typed, tested solvers and libraries (bmssp, min_ratio_cycle, heavytails, dynamical_systems_econometrics).
  • Reusable infrastructure: built developer tooling (smart-todo-action, git-actions-collection, article-reminders) to keep a large, multi-repo research output reproducible and maintainable.

Contribution activity across public and private work:

committers.top badge


Let's Connect and Collaborate

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors