Skip to content

About

Self-correcting MATLAB→Python translator. A rule-based transpiler does the deterministic pass; an LLM agent runs the output, checks the numbers, and patches until correct.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

MATLAB → Python Agent

A self-correcting agent that converts MATLAB code to working Python. A rule-based transpiler does the deterministic pass; an LLM agent then executes the result, inspects the output, and patches it until the numbers are correct.

The premise: rule-based transpilation has a hard ceiling, and the code it produces is often syntactically valid and semantically wrong. Running the code is the only reliable way to find that out.


Architecture

MATLAB source
     │
     ▼
┌──────────────┐
│  transpile   │  rule-based, deterministic, fast — handles the bulk
└──────┬───────┘
       ▼
┌──────────────┐
│  run_python  │  execute in a subprocess, capture stdout/stderr
└──────┬───────┘
       ▼
   correct? ──no──► agent patches the code ──┐
       │                                     │
      yes                                    └──► back to run_python
       ▼
  final Python

The agent is instructed to patch, not rewrite — preserving the transpiler's structural work and fixing only what is broken. Every rule fixed deterministically reduces the number of LLM turns needed; in one measured case, fixing a single regex bug in the transpiler cut a conversion from four agent steps to two.


Why execution-based verification

Test case 1 — the transpiler emits this, and it runs without error:

def sq(x):
    y = x**2

Valid Python. Exit code 0. Silently returns None for every input. Any check based on "does it run" marks this as a pass.

The agent instead ran print(sq(3)), saw None where 9 was expected, added the missing return, and re-verified with both a scalar and an array.


Results

Input Transpiler output Agent outcome
sq — element-wise square Valid Python, silently returns None Added return; verified on scalar and array
rowsum — loop with length() Invalid syntax: for i in range(1-1, len):(A) Repaired loop and indexing; verified sum against an independently computed expected value
normalize_cols — column normalisation Invalid and meaningless: np.zeros((size[A])), B[None:, j] Full NumPy rewrite using A.shape and [:, j]; verified the output matrix element-wise

On the third case the agent also added dtype=float unprompted — without it, integer array division would have truncated the results. It then constructed an expected matrix independently and compared, reporting Match? True.

It reached that in two steps.


Tools

Tool Purpose
transpile Rule-based MATLAB→Python conversion (separate project, consumed as a dependency)
run_python Executes code in a subprocess with a timeout; returns exit code, stdout, stderr

The system prompt encodes the transpiler's known weaknesses as a checklist — missing return statements, 1- vs 0-indexing, column-major linear indexing, struct-vs-method ambiguity, end in slices. This is a more effective use of domain knowledge than attempting to encode it in additional regex rules.


Engineering notes

Deterministic first, model second. Rules handle the predictable 80% at zero cost and zero latency. The model handles the long tail. Improvements to either half compound.

Running is not correctness. The agent is explicitly instructed to verify computed values, not exit codes, and to append its own test calls. In practice it generates its own expected values and asserts against them.

Subprocess isolation. Generated code runs via subprocess rather than exec(), so a crash or sys.exit() cannot take down the host process. A timeout catches the common failure mode of translated loop code running forever.

Bounded loop. Capped step count prevents a repeatedly-failing patch cycle from consuming quota indefinitely.


Security: why code execution is not exposed publicly

run_python executes model-generated code. The model composes that code from user-supplied input, which means a user who can steer the model can run arbitrary code on the host — reading environment variables (including API keys), making outbound network requests, or consuming resources. This is the "lethal trifecta": untrusted input, private data, and an external communication channel in the same system.

Input validation does not solve this. A check performed by the model cannot defend against an attack that manipulates the model. Real MATLAB can also contain system() and eval() calls that translate into genuinely dangerous Python, so "is this really MATLAB?" is not a safety property.

Accordingly:

  • The execution loop runs locally only — which is where researchers would use it anyway, on their own code and their own machine.
  • A public deployment would ship without run_python: transpile plus static model review, no execution. Less capable, zero remote-code-execution surface.
  • Proper public execution would require per-request container isolation — --network=none, read-only filesystem, no inherited environment, memory and CPU caps, non-root user, hard timeout — or a managed sandbox service.

Known limitations

  • Scope: numeric functions. No plotting, toolboxes, file I/O, or object-oriented MATLAB.
  • Verification is heuristic. The agent writes its own test cases; it does not compare against real MATLAB execution. Differential testing against Octave would be the rigorous version.
  • Silent semantic divergence remains possible for cases the agent does not think to test — column-major ordering, integer division, and broadcasting differences are the usual suspects.
  • Cost scales with transpiler quality. Poor rule coverage means more agent turns, which means more latency and more tokens.

Running locally

pip install -r requirements.txt
export GEMINI_API_KEY=your_key
python -c "from agent import convert; print(convert(open('example.m').read()))"

Stack

Python · Gemini API (function calling) · NumPy · subprocess isolation

License

MIT

About

Self-correcting MATLAB→Python translator. A rule-based transpiler does the deterministic pass; an LLM agent runs the output, checks the numbers, and patches until correct.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages