Written so this project can be picked up in a fresh session (or by a different person) without losing anything that currently lives only in conversation. Everything here is either not recorded elsewhere in the repo, or is a pointer to where it is.
Read order for someone starting cold: README.md → PORTING_PLAN.md
→ this file → docs/formats/.
Phases 0 and 1 are complete and Phase 2's core has landed. 83 tests pass; nothing is skipped and nothing xfails.
The Phase 1 milestone is met: python -m triton.tools.mkltsa produces .ltsa files that are
byte-for-byte identical to MATLAB's calc_ltsa output on all three fixtures — header, directory,
spare-slot zero fill, and every quantised spectrum. That is the unambiguous proof the plan asked
for, and it is the only test that covers the parameter derivation; every other test reads those
values out of a header rather than computing them.
| Phase | Status |
|---|---|
| 0 — spec + golden fixtures | Done. 12 fixtures, 190 reference artefacts, tests passing |
1 — core library (io, timing, dsp) |
Complete. 41 passing, 0 skipped, 0 xfail. Milestone met |
| 2 — session layer | Core landed. Milestone met: open, seek, spectrogram tile and LTSA tile, headless |
| 3 — GUI | Not started |
| 4 — plugin API + first Remora | Not started |
| 5 — tools | Not started |
| 6 — migration | Not started |
What exists under src/triton/:
| Module | Ports | State |
|---|---|---|
timebase.py |
the 2000-year datenum shift, wavname2dnum.m |
Done. All 6 filename patterns parse |
io/xwav.py |
rdxwavhd.m |
Done, including v2, which MATLAB's io/ readers cannot do |
io/audio.py |
readseg.m, check_time.m |
Done. Byte-exact against all 27 reference cases |
dsp.py |
hanning, mkspecgram.m, pwelch + int8, TF interpolation |
Done. Agreement is machine precision — see below |
io/ltsa.py |
read_ltsahead.m, read_ltsadata.m |
Done. Data blocks byte-identical on all 3 fixtures |
tools/mkltsa.py |
get_headers, ck_ltsaparams, write_ltsahead, calc_ltsa |
Done. Whole-file byte-identical output |
session.py |
the PARAMS global, control.m's update plumbing |
Core done. 42 tests in tests/test_session.py |
Nothing is skipped and nothing xfails.
Four MATLAB behaviours had to be reproduced rather than fixed, because existing LTSAs were
produced with them and a reader has to agree with the writer. All four are marked
BUG-COMPATIBLE at their sites in tools/mkltsa.py.
- The last spectral average of a raw file re-reads earlier data.
calc_ltsa.m:99advances the input pointer by this average's sample count, but the distance to move is the previous average's length. Identical for every average except the last one of a raw file that does not divide evenly, where the short length under-advances and the final average overlaps the one before it. It never reads outside the raw file, so nothing looks corrupt — the last time bin of each raw file simply is not the data it claims to be. This is live in the archive: the 10 kHz fixtures hit it (nave3 against an exact 2.5), so any LTSA whosetavedoes not divide its raw files evenly has this in every raw file's last bin. Worth deciding separately whether to fix it in MATLAB; doing so would change LTSA bytes and so needs its own conversation. fscomes from thefmtchunk, not the raw-file table (ltsa.md 5.1,ck_ltsaparams.m:20), and creation then aborts if any raw file disagrees. Thefakefsfixture is refused here and read happily by the display path — which is correct, and the error message says so.- 100 spare directory slots are reserved (
write_ltsahead.m:41), so a 612-byte-of-spectra LTSA is 11,492 bytes. Presumably for appending, which nothing does. - The filename field is padded twice — space-padded to the longest input name by MATLAB's char
matrix, then NUL-padded to 80 (
get_headers.m:113). Only observable with inputs of differing name lengths, which no fixture has; see the coverage note below.
Two test-side lessons, both the same shape as the ch offset bug:
test_python_mkltsa_is_byte_identicalwas passingr["fnames"]straight into the input list. That field is dumped per raw file, so a single four-raw-file x.wav appeared four times and the test built an LTSA of four copies of the same file. It now deduplicates, and reports the first differing byte offset on failure instead of dumping two 11 kB blobs — the offset alone usually names the field.- Every derived value matched on the first run once the input list was right. The failure was entirely in the harness, which is worth remembering: when a byte-identical test fails at a low offset, suspect the inputs before the arithmetic.
Measured agreement for dsp, because "it passes" understates it and a future regression should
have a number to fall short of:
| quantity | agreement | test tolerance |
|---|---|---|
hanning |
6e-16 absolute | 1e-15 |
| spectrogram dB | 7.3e-12 dB absolute, 3.2e-13 relative | 1e-9 |
spectrogram f, t |
exact (0.0) | 1e-12 |
pwelch dB |
9e-16 relative | 1e-9 |
| int8 quantisation | exact, all 3 cases | exact |
| TF interpolation | 2e-16 relative | 1e-12 |
Two things about dsp that took measuring rather than reading:
- State the spectrogram tolerance in absolute dB, not relative. A dB value passes through zero
legitimately —
specgram_nfft1000_ol50has a bin at −3.25e−4 dB — and a pure relative bound there measures its own noise. That one bin produced a 2.2e−10 relative error while its absolute error was 7.3e−14 dB, making the test look 4.5× from failing when it was eleven orders off. The assertion now carries both a relative bound and an absolute floor. - No pwelch fixture lands on a rounding tie, so the one part of
to_int8_dbthat needed custom code was passing untested. Closest approach across all three cases is 0.0054. MATLAB'sint8()rounds half away from zero; numpy'sround()rounds half to even, so they disagree on 0.5, 2.5 and 126.5. There is now a probe vector indump_reference.mand a test that checks the rule against documented behaviour immediately and upgrades itself to the MATLAB dump once one exists.
All seven recommendations at the end of PORTING_PLAN.md were reviewed and accepted on 2026-08-19, with these amendments from that conversation:
| # | Decision | Amendment |
|---|---|---|
| 1 | Integer samples + exact segment start times; datenum only at boundaries | Confirmed as the direction the group was already moving. If any MATLAB remains, move it to datetime, which does not degrade with epoch distance |
| 2 | Single TritonSession, one current_session() accessor, inspector panel + embedded console |
Accepted as-is |
| 3 | PySide6 + pyqtgraph, matplotlib for export | Accepted as a starting point, explicitly revisitable if a limitation shows up |
| 4 | Splice-and-delimit by default; honouring gaps as an option | Accepted; gap display was called out as a genuinely useful new feature, not just a compatibility switch |
| 5 | First-class Annotation model |
Accepted, with a wider mandate — see §3 |
| 6 | Entry-point plugin discovery, typed hooks, triton.api as the only import surface |
Accepted, with a low-barrier-to-entry constraint — see §3 |
| 7 | Don't port the dead files | Accepted enthusiastically; housecleaning is wanted, not risky |
These are not in the plan document and would be lost otherwise.
Annotation is a strategic capability, not a compatibility feature. It was central when Triton
was used for large-scale manual annotation and is underused now. The revamp should lay groundwork
for: dragging boxes rather than only xy picks; zoom; saving selected acoustic events for
classifier training or similarity search ("find me more like this"); and computing standard
metrics on a selection (amplitude variants, SNR). The Annotation dataclass in
PORTING_PLAN.md §4.4 should be checked against those use cases before it is
frozen — in particular it needs a time and frequency extent (it has one), a stable id, and
somewhere to hang derived metrics.
Remoras are written by students and junior research staff. Low barrier to entry is a hard requirement, not a nice-to-have: adding a plugin must stay easy and the instructions must be followable by someone who is not a software engineer. Balance this against the API in PORTING_PLAN.md §4.3, which is richer than what most plugins need.
The resolution to aim for: most existing Remoras are only nominally integrated — they open
from a Triton dropdown and are standalone thereafter. That path must stay trivial (one entry
point, one callable, done). The typed hooks (overlay, on_pick, on_segment_loaded) should be
strictly opt-in for the minority who want real integration. Design the tutorial around the
trivial case and document the hooks as a second tier.
Long-term goal is branch unification. Several divergent Triton versions exist across the group. The user wants everyone back on one main tool carrying everyone's features, with minimal regressions from version mixing. The Python port is the vehicle. This makes §4 below more urgent than it would otherwise be.
Measured 2026-08-19: the two trees produce byte-identical output on every path Phase 0 covers.
dump_referencewas run against bothTriton-masterandTriton_remorasand the two dumps compared: 74/74 binary artefacts and every JSON identical. Header parsing,readsegbyte arithmetic, derived timestamps,mkspecgramandpwelchare unaffected by the 19-file divergence.Not covered by that comparison, and still genuinely open: the LTSA generation path (
calc_ltsa,write_ltsahead,ck_ltsaparams), because the.ltsafixtures do not exist yet (§6.1). The knowncalc_ltsadifference affects only the final short time bin of wav/flac input, so even once fixtures exist the x.wav path may well come out identical too.Revised priority: this is not a Phase 1 blocker. Reconcile it as the port pulls on each file, using the harness as the arbiter (recipe below). The earlier "resolve before Phase 1" framing was written before the measurement and was too strong.
This is the useful by-product of Phase 0 for the consolidation effort: "which version is right?" becomes a measurement instead of a diff-reading exercise. About three minutes per tree.
matlab -batch "addpath('D:/Code/<CANDIDATE_TREE>'); addpath('D:/Code/Triton_python/tools/matlab'); \
dump_reference('triton','D:/Code/<CANDIDATE_TREE>','out','D:/Code/Triton_python/fixtures/_compare')"
# then diff fixtures/_compare against fixtures/reference; ignore _environment.jsonAny file whose differences do not move a single byte of that dump is a cosmetic merge — take either version. Files that do move bytes are the short list that deserves human judgement, and each one should gain a fixture case before the merge is finalised.
All Phase 0 specs were written against D:\Code\Triton-master. There is at least one other
full checkout on the same machine, and MATLAB's saved path actually points at the other one:
| Path | What it is |
|---|---|
D:\Code\Triton-master |
What the specs and the reference dump were built from |
D:\Code\Triton_remoras |
A second full checkout — same 90 base .m filenames, 19 differ in content. This is what the MATLAB path pointed to when dump_reference ran (visible as Triton_remoras\... warnings in its output) |
D:\Code\Triton\triton1.95.20230315 |
A versioned release drop |
Ignoring line endings, the 19 differing base files are:
calc_ltsa check_path ck_ltsaparams decimatewav decimatewav_dir
filepd get_headers get_ltsadir get_ltsaparams initdata
initpulldowns plot_ltsa plot_triton read_ltsadata read_ltsahead
triton wavname2dnum write_ltsahead xml_read
That is most of the LTSA pipeline. Neither tree is a superset of the other. Two confirmed examples, both spot-checked:
write_ltsahead.m— Triton-master has the un-scriptable~exist('PARAMS.ltsa.outfile','var')guard; Triton_remoras already has the fix (~isfield(PARAMS.ltsa, 'outfile')), added for the BatchLTSA remora, and identical to the fix independently proposed here.calc_ltsa.m— Triton_remoras branches on vector orientation when zero-padding the short final time bin; Triton-master does not, and is wrong for wav/flac column-vector input. This changes numeric output, so archived.ltsafiles may contain either behaviour.- Going the other way, Triton-master has flac support (
ftype == 3) that Triton_remoras lacks.
Action: not a gate on Phase 1. Two things worth doing early and cheaply:
- Get every uncommitted tree into version control now. Pure risk reduction, no scope creep,
and
Triton_remorashas already been shown to hold the authoritative version of at least two files. Losing it would lose real work. - Build the
.ltsafixtures (§6.1), which closes the one measurement gap above and settles thecalc_ltsaquestion with data rather than argument.
Everything else — deciding a canonical version per file — can be pulled by the port, one file at a time, arbitrated by the recipe above. Details in ltsa.md §7.
Established in conversation, now also recorded in xwav.md:
Files carrying
% Do not modify the following line, maintained by CVS/$Id: ... $were written by a different developer from the core reader/writer set. They often re-express the same logic more clearly but are re-implementations and can carry errors. The un-marked core files (rdxwavhd,wrxwavhd,readseg,read_ltsahead, …) were written by the author of the XWAV format and are the authority when in doubt.
60 files in Triton-master carry the CVS marker, including all of io/ and prefix.m in the base
folder. This rule already decided one case: the v2 header bug in §6 below is in a CVS-marked file,
so rdxwavhd wins.
Milestone met: tests/test_session.py::test_milestone_open_seek_and_retrieve_both_tiles opens a
file, seeks, and returns both a spectrogram tile and an LTSA tile with no GUI toolkit imported.
- Position is an integer sample index, never a float time.
AudioStateholds(segment, offset);timeis a derived property. This is the design rule from timebase.md applied to application state, and it removes the path-dependence documented inAudioSource.skip_for—step(+1)thenstep(-1)returns to the identical sample, and ten thousand steps accumulate no error. Two tests pin that, including one with a step size that is not a whole number of samples. - Mutation is observable.
session.subscribe(cb, prefix=...)gets dotted paths like"view.nfft";session.batch()collapses a group of changes into one round so a subscriber never redraws an intermediate state nobody asked for. This is what replacescontrol.m's ~50 hand-rolled update sequences. - Validation happens at assignment.
view.nfft = -1raisesValidationErrorthere, rather than surfacing inside a plot routine three calls later, which is where MATLAB finds these. - No GUI toolkit is imported, enforced two ways — a
sys.modulescheck and an AST scan ofsession.py's own imports. Thesys.moduleshalf is weak on a machine without Qt installed; the AST scan is always meaningful. Phase 3 adapts the callback registry to Qt, not the reverse. - The session adds no arithmetic. Four equivalence tests assert that a tile obtained through
the session is identical to the same data obtained by calling
triton.ioandtriton.dspdirectly. Without those the session could quietly become a second implementation and the parity suite would not notice, because it never goes through the session.
_max_start_offset has two bounds and both are load-bearing. The furthest a window may start
in a segment is limited by (a) the last addressable sample, which is not the last sample
present — a segment's declared end is one sample short of its data (OPEN_DECISIONS §1.5) and
segment_containing tests t < end strictly, so the final sample resolves to no segment at all;
and (b) in the last segment only, room for a full window, matching check_time.m:34-50's
"too late" rule. Earlier segments are deliberately not bounded by (b), because a window there is
meant to spill into the next raw file — that is the splice behaviour analysts rely on. Both
bounds were found by tests failing, not by reading.
Clamping is not reversible, and that is correct. Once a step clamps at either end, how far past
the end it went is gone, so a matching step back does not return to where it came from. The
round-trip exactness claim is therefore qualified: it holds while the position stays inside the
file. An earlier version of that test used 500 steps in a 100,000-sample file, clamped in both
directions, and passed while demonstrating clamping rather than exactness — it now computes a step
count that stays inside and asserts not at_end() so it cannot silently degrade again.
as_params() is a fresh dict every call, deliberately. Handing out a live handle would
reintroduce exactly the thing PARAMS is criticised for: the complaint was never that it was
inspectable, it was that forty functions could mutate it. A test asserts that mutating the returned
dict does not reach the session. Per the 2026-08-23 decision it is also free to be clearer than
MATLAB rather than bug-compatible — it exposes per-raw-file dt (which rdxwavhd.m overwrites)
and real instants rather than shifted datenums, and says so in its docstring.
Undo (the plan lists it as something the design buys, and nothing needs it yet), the Session
Inspector panel (walk() is the data source it will use), the embedded console, and the plugin API
proper — that is Phase 4, and session.plugins is only the namespace it will use.
Four things turned up while implementing, each of which changes something outside the port.
Two writers disagree by one character, and both kinds of file are in the archive:
| writer | code | byte written |
|---|---|---|
HARPproc/write_XWAVhead.m:42 |
PARAMS.xhd.WavVersionNumber = 1; |
0x01 |
HARPproc/mk_SpotCheck.m:391 |
PARAMS.xhd.WavVersionNumber = '1'; |
0x31, ASCII '1' |
Both reach wrxwavhd.m:145, which writes with fwrite(..., 'uchar'); that converts a char to its
code point. In mk_SpotCheck.m every neighbouring assignment (InstrumentID = 'DLXX', SiteName = 'sitX', DiskSerialNumber = '12345678') genuinely is a string, so the quotes read as
copy-paste consistency rather than intent.
Measured over Triton_remoras/ExampleData — 123 x.wav files, all of them real:
| version byte | firmware | instrument | files |
|---|---|---|---|
1 |
V2.02S |
DL37 |
5 |
1 |
V2.87 |
D102 |
4 |
1 |
V2.98 |
D111 |
2 |
49 = '1' |
3A01240501 |
DLXX |
112 |
The split is by writer exactly as predicted; the 112 are mk_SpotCheck output, and their scrubbed
instrument/site/serial values are that function's literals rather than a separate anonymising step.
Latent, not cosmetic. v0 and v1 have identical layout and MATLAB's only consumer is
WavVersionNumber == 2, false either way, so nothing misreads today. But a v2 file written
through the mk_SpotCheck path would carry '2' = 50, == 2 would be false, rdxwavhd.m:106
would size the harp chunk with the v1 formula, and every byte_loc in the raw-file table would be
read from the wrong offset — after emitting only its generic "SubchunkSize and NumOfRawFiles
discrepancy" warning. This compounds the four Remora copies of io/ioReadXWAVHeader.m, which have
no v2 branch at all (xwav.md 6.8).
Two MATLAB-side actions, neither done:
- Drop the quotes in
mk_SpotCheck.m:391.mk_SpotCheck.mis in HARPproc, which is not merged intoTriton_remorasyet, so this belongs to that merge. It fixes new files only; the 112 existing ones keep their49. - Normalise the field in
rdxwavhd.mright after thefread, so the== 2tests work whichever way the byte was written. This is the change that protects existing and future data, and it is the one worth running through the MATLAB regression harness. No header has version 48 or higher, so mapping 48/49/50 to 0/1/2 is unambiguous.
io/xwav.py already accepts both. Full detail in xwav.md §3.
readseg.m:111 computes floor((plot.dnum - dnumStart) * 86400 * fs) from float datenums
carrying about 40 ns of representation error. Forty nanoseconds is a small fraction of a sample at
any rate Triton handles, so it changes the answer only when the requested time lands within that
distance of a sample boundary — and there it decides the floor. Of the 27 reference cases, 7
land exactly on a boundary, and MATLAB falls one sample low in all 7:
offset 0.625 s at 10 kHz -> exact 6250, MATLAB 6249 (product 6249.999889)
offset 0.37 s at 10 kHz -> exact 3700, MATLAB 3699 (product 3699.999812)
Reproducing it is not merely hard, it is ill-defined: plot.dnum is not recomputed from the
requested time, it is accumulated — the file's start plus however many forward and backward steps
the analyst took. Two routes to the same nominal time carry different accumulated error, so the
same time in the same file can read different samples depending on how the analyst got there.
There is no "the MATLAB answer" to match.
So the port computes exactly and may differ from any given MATLAB session by at most one sample, only at exact sample boundaries. That is 100 µs at 10 kHz and 5 µs at 200 kHz — immaterial for spectrograms and detection, and it is the correct sample for anything that cares, such as cross-correlation time-of-arrival work.
This shapes the test design, and the shape matters more than the finding. read_samples(index, skip, n) is public API precisely so the parity test can ask for exactly the samples a given
MATLAB session read; that comparison is then byte-exact with no tolerance at all. The one-sample
question is isolated in test_time_addressed_read_is_within_one_sample_of_matlab, which asserts
both halves — never more than one sample out, and never out at all unless the exact product really
is within a datenum quantum of an integer. A genuine indexing bug fails the second half
immediately. Do not collapse these back into one fuzzy-tolerance test.
XwavSource deliberately keeps both. sample_rate is the raw-file table's, which is what
skip uses (rdxwavhd.m:188, whose own comment says the fmt value "could be fake").
byte_rate is the fmt chunk's, which is what raw-file end times and delimiter spacing divide
by (rdxwavhd.m:160, readseg.m:129). On the fakefs fixture these are genuinely different
numbers and using one for the other shifts every delimiter. Both are reproduced as written.
test_ltsa_header_matches originally asserted six of the fourteen fields the reference dumps.
ch was not among them, and the LTSA reader read it from the wrong offset — v3/v4 store it at 38
and the reader looked at 39. Offset 39 is padding, padding is zero, and zero is a perfectly
plausible channel number, so the only symptom was a quietly wrong value on every v4 file. The
test passed.
It surfaced only because a hand-written sanity script printed the field next to the value the
fixture was built with (ch = 1) and they disagreed. Two changes followed:
- the test now asserts every scalar the reference dumps, plus
rfileidandfnamesper entry, plus that thebyte_locchain reproduces fromnave × nf(ltsa.md 3.2) — which is checkable without MATLAB and would catch a wrong directory stride; io/ltsa.pyhas a_check_layout()that runs at import and asserts the header and entry geometry of all four versions sums to its declared size, so an offset edit that breaks the arithmetic fails loudly instead of returning zeros.
The general form is worth keeping in mind for the remaining phases: a dumped reference field that no assertion touches is not coverage. Prefer asserting the whole record.
The probe added to dump_reference.m has been run. MATLAB's int8() output on the 15 probe values
matches to_int8_db exactly, and numpy's default round() would have differed on 4 of them
(−0.5, 0.5, 2.5, 126.5). That gap is closed with evidence.
check_time.m leaves PARAMS.raw.currentIndex empty when the time is past the last raw file, and
of length 2 when it falls inside two at once. readseg.m:108 then indexes byte_loc with it and
hands fseek a non-scalar offset, which errors out with nothing that names the cause. The comments
present both as warnings; they are not. The port raises TimeNotInData / AmbiguousTime from the
place that knows why, and the message says which raw files and what to do.
-
LTSA fixtures are not built.Done, 2026-08-23. The first of the three options worked: pointedmake_ltsa_fixtureatTriton_remoras, whosewrite_ltsaheadalready tests~isfield(PARAMS.ltsa,'outfile')rather than the brokenexist('PARAMS.ltsa.outfile','var'), and it ran headless with no dialogs. Three v4 LTSAs built,dump_reference('sections', {'ltsa','dsp'})captured, all four previously-skipped tests now pass.Two things had to be fixed in
dump_reference.mfirst, both consequences of the MATLAB side having improved since Phase 0 was written:- Its
ltsasection hand-buildsPARAMSand calledread_ltsadatadirectly. That worked whencheck_ltsa_timewas commented out atread_ltsadata.m:11— which is what issue #129 turned out to be — and threwUnrecognized field name "step"once it was restored. It now setstseg.step,plotStartRawIndexandplotStartBin, which is whatinit_ltsadatawould have contributed. Note the ordering trap:read_ltsadatadoes deriveplotStartRawIndexitself, but only aftercheck_ltsa_timehas already read it, so omitting it is an immediate error rather than a latent one. - No real caller was affected —
initparams.m:131,mk_ltsa.m:50andsm_mk_ltsa.m:46all settseg.step— so this was fixture tooling, not a regression. Checked before assuming.
Then
dump_reference('sections',{'ltsa'}). Until this happens,test_ltsa_reference_present_or_explainedand the LTSA parity tests skip. - Its
-
Real deployment data.
fixtures/real/is no longer empty — it now holdsGOM_AC_05_disk01_212287.x.wav(30 MB), its 15 MB raw counterpart, andGOM_AC_01_example.ltsa. Correctly gitignored; only.gitkeepandREADME.mdare tracked.It earned its keep immediately: the
WavVersionNumberfinding in §5a came out of the first real file, which the reader rejected outright. Synthetic fixtures prove we implemented the spec; real files prove we implemented reality, and here the spec was wrong.The x.wav now reads end to end — 200 kHz, 1 channel, 16-bit, one raw file, 75 s, and
byte_loc + byte_lengthequals the file size exactly. Two follow-ups:fixtures/real/MANIFEST.mdstill says "(none yet)". Worth filling in, and note that.gitignorewhitelists onlyREADME.mdunderfixtures/real/, so the manifest is currently untracked. Add!fixtures/real/MANIFEST.mdif it should be committed — the point of a manifest is that it outlives the data.- The remaining wish list in fixtures/README.md still stands, and one
entry is now sharper: a v2-header file would settle both the
== 2question in §5a and the severity of theio/reader bug.
The Phase 1 io tests pass, but three behaviours have no MATLAB reference to check against. They are implemented from the source and hand-verified; that is weaker than the rest of the suite, and worth closing when MATLAB is next in front of someone.
-
The
check_time.mtriage is untested. All 27readseg_indexreference cases havecurrentIndex == 1, so conditions 2a (a time in a duty-cycle gap, and the direction-dependent snap), 2b (past the last raw file) and 3 (inside two raw files at once) are exercised by nothing. Closing it needs three new fixture cases: an offset landing inside a gap onxwav_v1_duty_1ch_16b_10k.x.wav, an offset past the end, and a file written with overlapping raw files. The gap case is the one that matters — it is how analysts page through duty-cycled data. Hand-verified meanwhile: the backward snap lands atend - (tseg - 1/fs), which for raw file 0 of the duty fixture is08:45:02.4999 - 0.4999 = 08:45:02.000000000exactly. -
Gain on a multichannel file is untested. No fixture has
nch > 1andgain > 0, so the quirk inreadseg.m:117-119— the division is applied to columnPARAMS.chalone, leaving the other channels unscaled — is reproduced but unverified. If any real 4-channel deployment ran with gain, its other three channels have always come back unscaled. -
splice_gaps=Falsehas no oracle by construction. No MATLAB path lays a window out on true wall-clock time, so the NaN-filled mode can only be checked for self-consistency. It is: on the duty fixture a 0.5 s window opened at +2.25 s gives 2500 real samples then 2500 NaN, the switch falling exactly where raw file 0's 2.5 s of data ends, and the real samples are identical to the spliced mode's. -
The writer is only checked on single-input LTSAs. All three fixtures are one x.wav each, so two things are unexercised: the filename space-padding described above (which needs inputs of differing name lengths), and the
bytelocchain across a file boundary as opposed to across raw files within one file. Both are cheap to close — build a fourth fixture from two of the existing x.wavs, whose names differ in length, and re-dump. Worth doing before anyone relies on multi-file output. -
split_inputshas no end-to-end test. The chunking arithmetic is trivial and tested by inspection, but nothing builds two LTSAs from one oversized list. It cannot be tested honestly without 65,536 input files, so a smallermax_filesshould be injected instead.
The 24-bit decode path is also unverified, but that is expected and harmless — no 24-bit x.wav exists (§6, Resolved).
Everything in this subsection, plus the deferred items below and the format-doc "known inconsistencies" lists, is now collected in docs/OPEN_DECISIONS.md with the evidence, what would change if each were fixed, and who can settle it. That file is the canonical register; this table is kept as a quick index. Add new items there.
| Question | Status |
|---|---|
32-bit AudioFormat: wrxwavhd writes 3 (float), readseg reads int32 |
Open. Needs one real 32-bit x.wav. Fixture currently writes AudioFormat = 1 with int32 samples |
blksz hard-codes the 16-bit assumption (ck_ltsaparams.m:50), so write_length bookkeeping is questionable for 32-bit |
Open, coupled to the above |
| Do x.wav files with channel counts other than 1 or 4 exist? Triton's sector geometry only defines those two | Open |
| Do v2-header x.wav files exist in the archive, and in what quantity? | Open, and it determines the severity of the io/ bug below |
| Which tree is canonical (§4) | Open, highest priority |
- 24-bit x.wav.
readseg.masksfreadfor a nonexistent'int24', so it cannot read them. Confirmed with the team: no 24-bit x.wav exist; the archive is all 16-bit. Implement it in Python anyway (nearly free) but it needs no fixture and blocks nothing. 24-bit plain wav does occur and has a fixture.
rdxwavhdper-raw-filedt/paddingoverwrite. Lines 135/137 assign unsubscripted inside the loop, so only the last raw file's values survive. This may be intentional, and some Remoras are believed to compensate for it. Plan: the Python reader exposes the full per-raw-filedtarray (strictly more information, so nothing breaks by reading it), but no change to what consumers see until the compensating Remora code is found and audited. Do that audit during Phase 4, not before.io/ioReadXWAVHeader.mhas no v2 branch — it skipsdrate/dtand returns wrongbyte_loc/byte_lengthsilently. Copies live in four Remoras (Detector/io/,SPICE-Detector/io/,BlueWhaleBcall-Detector/Detection/, rootio/). Agreed to fix when Remora work begins, not now. If v2 files turn out to be common, results those Remoras produced from them are suspect and this becomes urgent. Per §5,rdxwavhdis the correct reference.write_ltsaheadscriptability — agreed to make LTSA creation scriptable in the future. The fix already exists inTriton_remoras; folding it in is part of §4.
- MATLAB R2025b at
C:\Program Files\MATLAB\R2025b\bin\matlab.exe. R2013b, R2016b, R2023a and R2024a are also installed. The reference dump records which release produced it infixtures/reference/_environment.json— numerics can shift between releases. - Headless dump command that works:
Takes a couple of minutes, mostly MATLAB startup.
matlab -batch "addpath('D:/Code/Triton-master'); addpath('D:/Code/Triton_python/tools/matlab'); dump_reference"dump_reference('sections',{...})runs a subset. tr_headless_handles.mis the trick that makes the oracle trustworthy. It builds the ~15HANDLESthat core functions touch, soreadseg,check_time,mkspecgramandread_ltsadatarun for real headlessly instead of being re-implemented in the dump script. If a new dump section needs another Triton function, extend that shim rather than transcribing logic.- Use the project venv, not system Python. System Python 3.12 has a broken
anyio/pytest plugin (missingtyping_extensions) that makes anypytestinvocation fail with an unrelated traceback..venv/(numpy 2.5.1, scipy 1.18.0, pytest 9.1.1) is clean. - Fixture generation uses
numpy.random.RandomState, notdefault_rng— NEP 19 guarantees the former's stream is stable across numpy versions forever. Verified byte-identical output under numpy 1.26.2 and 2.5.1. Do not "modernise" this. - Committed data sizes:
fixtures/generated/2.3 MB,fixtures/reference/3.7 MB. The reference dump storesreadsegblocks as int32 where lossless and caps windows at 25k samples to keep that number down; if you add dump sections, keep an eye on it. - The repo is not a git repository yet.
git initand an initial commit is a reasonable first action in the new environment.
Done since the last revision: git init; triton.timebase; triton.io.xwav;
triton.io.audio (readseg + check_time). Remaining, in order:
- Phase 3 — GUI v1, the long pole. Or, if a smaller next step is wanted, the pieces of
Phase 2 listed as not-built-yet above: the Session Inspector is the cheapest of them and
walk()already supplies its data. - Close the two writer coverage gaps above — a two-input LTSA fixture with names of differing length. Small, and it is the one place multi-file output is currently unproven.
- Resolve §4 — pick the canonical MATLAB tree, or merge the two. Re-run
dump_referenceafterwards if anything in the pipeline changed. - Close the parity coverage gaps in §6 while MATLAB is open anyway; the duty-cycle gap case is the one worth the effort.
- Then the session layer (Phase 2).
- Phase 1 milestone:
test_python_mkltsa_is_byte_identicalgoes from xfail to pass.
Carried over to the MATLAB side, not this repo: the two WavVersionNumber fixes in §5a.
| File | Why |
|---|---|
| PORTING_PLAN.md | Architecture, phases, the PARAMS/session design, the plugin API sketch |
| docs/formats/xwav.md | Byte tables + 8 known MATLAB inconsistencies |
| docs/formats/ltsa.md | Byte tables + the two-tree divergence + 5 known issues |
| docs/formats/timebase.md | The year-2000 offset and why it must not be "cleaned up" |
| tests/test_parity_phase1.py | Phase 1's definition of done, already written |
| tools/matlab/dump_reference.m | How the oracle is produced |