PhD Candidate in Economics @ Universidad Autónoma de Madrid
Verified AI for Empirical Research
Causal inference · Agent evaluation · Climate & policy applications
I work at the intersection of reliable AI, causal inference, and empirical policy research. I build execution-grounded AI agents and benchmarks for research workflows where conclusions should be supported by evidence, code, and reproducible execution — not just plausible text.
Core question: Can AI agents complete reliable empirical research workflows, and can we verify when their scientific conclusions are actually supported by data and execution?
| Direction | Focus |
|---|---|
| Verified AI & Agent Evaluation | Execution-grounded benchmarks, tool use, evaluator validity, robustness, reproducibility |
| Causal Inference | Difference-in-differences, event studies, policy evaluation, causal workflow auditing |
| Climate, ESG & Policy | Carbon markets, climate policy, ESG measurement, auditable AI for empirical research |
An execution-grounded benchmark for evaluating reliable LLM tool use under adversarial and unreliable trade-data API conditions.
A multilingual ESG-perception research project built around an auditable agentic LLM pipeline. The research repository is currently private while the project is under development.
A document-grounded agent for retrieval and reasoning over U.S. Treasury Bulletin data.
A public research demo for structured and auditable classification of climate-related claims.
A one-command LLM-powered research knowledge-base workflow for Claude Code + Obsidian.
- Green Comtrade Bench — deterministic benchmark and judge infrastructure supporting the ComtradeBench research line.
- Purple Comtrade Baseline — deterministic reference agent used to validate the benchmark contract.
- AgentBeats Leaderboard — submission and evaluation infrastructure for the earlier AgentBeats benchmark configuration.
I care about whether an empirical or AI-assisted research workflow can be:
- executed against real tools and data
- checked against explicit assumptions and evidence
- reproduced by another researcher
- audited when something goes wrong
- interpreted correctly before scientific conclusions are trusted
Research identity: Verified AI for empirical research — causal inference, agent evaluation, and climate/policy applications.

