Skip to content
View yonghongzhang-io's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report yonghongzhang-io

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
yonghongzhang-io/README.md

Yonghong Zhang

PhD Candidate in Economics @ Universidad Autónoma de Madrid
Verified AI for Empirical Research

Causal inference · Agent evaluation · Climate & policy applications

I work at the intersection of reliable AI, causal inference, and empirical policy research. I build execution-grounded AI agents and benchmarks for research workflows where conclusions should be supported by evidence, code, and reproducible execution — not just plausible text.

Core question: Can AI agents complete reliable empirical research workflows, and can we verify when their scientific conclusions are actually supported by data and execution?

Research

Direction Focus
Verified AI & Agent Evaluation Execution-grounded benchmarks, tool use, evaluator validity, robustness, reproducibility
Causal Inference Difference-in-differences, event studies, policy evaluation, causal workflow auditing
Climate, ESG & Policy Carbon markets, climate policy, ESG measurement, auditable AI for empirical research

Featured Projects

An execution-grounded benchmark for evaluating reliable LLM tool use under adversarial and unreliable trade-data API conditions.

🌍 GLEAM (work in progress)

A multilingual ESG-perception research project built around an auditable agentic LLM pipeline. The research repository is currently private while the project is under development.

A document-grounded agent for retrieval and reasoning over U.S. Treasury Bulletin data.

A public research demo for structured and auditable classification of climate-related claims.

A one-command LLM-powered research knowledge-base workflow for Claude Code + Obsidian.

Supporting Benchmark Infrastructure

  • Green Comtrade Bench — deterministic benchmark and judge infrastructure supporting the ComtradeBench research line.
  • Purple Comtrade Baseline — deterministic reference agent used to validate the benchmark contract.
  • AgentBeats Leaderboard — submission and evaluation infrastructure for the earlier AgentBeats benchmark configuration.

Research Principles

I care about whether an empirical or AI-assisted research workflow can be:

  • executed against real tools and data
  • checked against explicit assumptions and evidence
  • reproduced by another researcher
  • audited when something goes wrong
  • interpreted correctly before scientific conclusions are trusted

Research identity: Verified AI for empirical research — causal inference, agent evaluation, and climate/policy applications.

Pinned Loading

  1. comtrade-openenv comtrade-openenv Public

    ComtradeBench OpenEnv: execution-grounded benchmark for reliable LLM tool-use under adversarial trade-data API conditions.

    Python 2

  2. purple-agent-officeqa purple-agent-officeqa Public

    A2A retrieval agent for Treasury Bulletin OfficeQA: document-grounded QA over public finance tables and reports.

    Python

  3. climate-claim-classifier climate-claim-classifier Public

    An LLM/NLP demo: classifying climate-related claims (type, carbon-market relevance, specificity, evidence needed)

  4. phd_kb_starter phd_kb_starter Public

    One-command installer for an LLM-powered personal knowledge base, following Andrej Karpathy's 8-stage workflow. Built for Claude Code + Obsidian.

    Shell 1 2

  5. agentbeats-leaderboard-v2 agentbeats-leaderboard-v2 Public

    Leaderboard infrastructure for the ComtradeBench / AgentBeats agent-evaluation benchmark: task definitions, submission flow, and scoring.

    Python 1

  6. green-comtrade-bench-v2 green-comtrade-bench-v2 Public

    Deterministic offline ComtradeBench judge for evaluating agent robustness under pagination, retries, duplicates, page drift, and totals traps.

    Python 1