Files
XZBT/reviews/.completed-artifacts/03-deepseek-v4-flash.md
T
LabyricornandClaude Opus 5 c4332363a9 docs: raise the format specification to revision 0.9 and land the reconciliation
Revision 0.9 adds section 20, the cadence and event subsystems contract, and
carries two corrections the implementation forced. Section 6.1 now states that a
duration is the authored literal or a non-negative finite number already in
milliseconds, since a DurationSpec may be the resolved output of a ValueSpec or
a bounded TimeSpec, with the one documented exception of an automation track's
`at`, which 19.1 keeps literal-only so that point ordering stays decidable at
import. Section 20.11 documents the rejection of an undeclared input name in an
event action's `with` map as ERR_UNKNOWN_FIELD — the section's own convention
for that shape of error, replacing an invented code that appeared nowhere in the
registry.

The review record is committed with the code it describes: the two code triages
that found these defects, the reconciliation plan that sequenced the fixes, and
a follow-up debt record listing what was deliberately left open — the unchecked
JSON Schema artifact, degenerate path arcs, post-effect transient allocation,
the window-traffic fixture's per-copy wrap bounds, and the unstated
`ownership: "persistent"` value on a sound action. None of the five blocks phase
6; all five are written down rather than dropped.

Devlog entries are backfilled for the two milestones that had none: phase 3c
slice 2, the audio lifecycle and voice ceilings, and slice 4d, the renderer
core. The implementation status summary now reflects the reconciled state rather
than the in-flight one.

231 tests pass. tools/verify-spec-contract.py reports 46 declared diagnostic
codes with every used code resolving and its two long-standing unresolved
cross-references unchanged.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01ShxxFqFmCUDQnQvFNm4TKy
2026-09-06 21:54:09 +00:00

17 KiB
Raw Blame History

XZBT format specification sections 14–16 — defect review

Reviewer: independent pass over docs/XZBT_0-1_Format_Specification.md (commit e5ed468), sections 14–16 only. Scope: defects only — no rewrites, no features, no style commentary. Line numbers refer to the spec file at that commit. Not reviewed against any other review file; this analysis is standalone.

Ranking: findings that make the contract unimplementable or untestable outrank inconsistencies and under-specifications.


F1 — No target supports both automation and override; 16.2's masking rule and acceptance trace 16.11.8 cannot be executed, and the bus-gain "automation" claim contradicts 16.1's own target rule

Sections: 16.1 (lines 1286, 1293), 16.2 (lines 1322–1332), 8.1 (line 314), 15.17 (line 1230), 14.4 (line 631), 16.11 trace 8 (line 1485).

What the document says:

  • 16.2 gives the resolution pipeline "For a node property: base ValueSpec -> binding (where supported) -> automation -> winning override -> modulation sum -> safety clamp -> engine parameter" and states as a consequence: "An override masks the automation stage's output for as long as it wins… When the override releases, the property returns to whatever the track has reached by then."
  • 16.11 trace 8: "An override masking an automated property releases to the track's current value, not the value held when the override took hold."
  • 8.1, 15.17, and 16.2 all state that audio.buses.<id>.gain "takes binding, automation, override, and modulation."

Why it is wrong: The two capability sets never intersect on any property a document can actually author:

  • Node properties support automation (16.1) but are barred from override: 14.4 — node fields are not added to the 8.1 table, and "an override action addressing a node field … is ERR_UNSUPPORTED_TARGET." 16.2 itself reaffirms this.
  • parameters.<id> and state.<id> support override but have Automation = No (8.1).
  • Bus gain is the only target with both Automation = Yes and Override = Yes (8.1), but 16.1 confines automation tracks to a graph object's automation array with target = <node-key>.<property>, "resolved in the same graph," restricted to "any property in the 15.13 modulatable registry, and no other." Buses are not nodes in any graph (they are top-level audio.buses.<id> objects whose only field is gain, 15.17), and bus gain is not in the 15.13 registry. No authoring surface for a bus-gain automation track exists anywhere in sections 14–16.

Consequences: (a) 16.2's masking rule describes a situation no legal document can produce, and (b) required acceptance trace 16.11.8 is unimplementable — the Phase 3c acceptance suite cannot pass as written. The same orphaning applies to bus-gain "additive modulation": the only modulation syntax in the document is graph-scoped routes of 15.12/15.13, whose to must be <node-key>.<property> in the same graph, so nothing can additively modulate a bus gain either, despite the 8.1 row and the 15.17/16.2 sentences.

Smallest fix: Pick one side and make the other consistent — either (a) delete "automation" (and, for the same reason, "additive modulation"/"modulation") from the bus-gain capability claims in 8.1, 15.17, and 16.2, and delete or re-scope 16.2's masking consequence and trace 8 to a target that actually exposes both stages — no such target exists, so the masking rule and its trace must be removed unless a target with both stages is added — or (b) give bus gain an actual automation authoring surface and make 16.1's target rule name it. Option (a) is the smaller change.


F2 — 16.6 eviction step 3 cannot free a budget slot; the outcome for the incoming instance is undefined

Section: 16.6 (lines 1406–1415); also 16.3 (line 1345) and 16.11 trace 10 (line 1487).

What the document says: "A voice counts against its ceiling from CREATED until DISPOSED." The eviction policy is: "1. Dispose the oldest instance already in FINISHED. 2. Evict the oldest instance in RELEASING by advancing its release ramp to immediate completion and disposing it. 3. For a one-shot request only: evict the oldest ACTIVE one-shot by starting its release. 4. Otherwise refuse the new instance." The runtime "stops at the first candidate," and "Eviction always releases (16.4); it never hard-stops an active voice."

Why it is wrong: Steps 1 and 2 dispose an instance, so a slot frees immediately. Step 3 only starts the release of an ACTIVE voice, moving it to RELEASING — and a RELEASING voice still counts against the ceiling until DISPOSED (that is precisely why step 2 must advance the ramp to completion and dispose). So when the ceiling is full of ACTIVE one-shots (no FINISHED, no RELEASING), step 3 frees nothing, yet the policy stops there. Admitting the new instance then puts the count one over the ceiling (65 > 64), contradicting the "counts from CREATED until DISPOSED" rule; refusing it contradicts the policy having stopped at a candidate rather than reaching step 4. Neither outcome is defined. Acceptance trace 16.11.10 (exercise the eviction order at the one-shot ceiling) cannot be implemented unambiguously in exactly this state.

Smallest fix: One sentence defining the step-3 outcome, e.g. that an evicted ACTIVE one-shot is removed from the ceiling count immediately upon eviction (its release then runs without occupying budget), so the new instance is admitted and the ceiling is never exceeded; or make step 3 behave like step 2 (complete the release immediately and dispose). Either way the "counts until DISPOSED" sentence needs the matching exception or wording.


F3 — 16.5's ending bound is not an upper bound when a bound-contributing property is modulated

Sections: 16.5 (lines 1376, 1385), 15.6 (lines 943, 949), 15.13 (line 1082), 16.3 (line 1360).

What the document says: "Determinable" means "the runtime can compute a finite upper bound on the instance's audible duration at instantiation, from the resolved graph alone, without observing output." The bound table gives delay the contribution time x ceil(log(1/1000) / log(feedback)). The instance is finished (torn down) at that bound plus the release duration.

Why it is wrong: delay.time is modulatable (15.13, in milliseconds) and is clamped live so it never leaves [0ms, 10s] (15.6). A legal oneshot can route an lfo or sample-hold to a delay's time; those sources vary over time and their future output is unknowable at instantiation. The 16.5 row computes the contribution from the resolved time only, but the actual delay time — and therefore the actual time the feedback tail takes to fall 60 dB — can be up to the 10 s clamp. The stated bound is then not an upper bound on audible duration: the runtime reaches the computed ending and tears the instance down while its tail is still ringing above the level the bound assumes, directly contradicting 16.3's claim that at the determinable ending "the envelope has already returned to zero." Trace 16.11.9's "bound matches the 16.5 table" only passes for graphs whose bound-contributing properties are not modulated.

Smallest fix: In 16.5, define the delay contribution over the largest delay time the property can reach given its modulation routes (source outputs are bounded — lfo by its resolved amplitude, sample-hold by its resolved min/max — and the property is clamped), or exclude a modulated bound-contributing property from determinable-ending status. One sentence either way; the first option preserves the most legal graphs.


F4 — 14.12's min/max ordering check fires after resolution but its code is filed Semantic-only: recurrence of the known staging defect

Sections: 14.12 (lines 812–819), 14.4 (line 629), §7 (line 286).

What the document says: min and max are ValueSpec<number> fields "resolved once, at the owning sound instance's instantiation boundary" (14.4). 14.12: "min resolving to a value greater than or equal to max is ERR_INVALID_RANGE_ORDER." Section 7 lists ERR_INVALID_RANGE_ORDER with stage Semantic and cause "Declared paired bounds (e.g. sample-hold min/max) are not in strictly increasing order after resolution."

Why it is wrong: This is the same defect class that has already occurred twice in this document (the 14.5 audioMaxFrequency semantic-stage issue, and the automation-ordering rule fixed in 50fb72c), and it is still present here. For literal bounds the check is import-decidable, but for procedural bounds the order exists only "after resolution," which happens at instantiation — no AudioContext and no instance exist at semantic validation, so a pair such as min: {random:{-1,1}}, max: {random:{-1,1}} that resolves inverted is undetectable at import, and the instantiation stage may not raise a code the section 7 table files as Semantic-only.

Smallest fix: File ERR_INVALID_RANGE_ORDER as "Semantic / Runtime" in section 7 (as ERR_OUT_OF_BOUNDS already is), with the cause text stating literals are rejected at import and resolved pairs at instantiation.


F5 — 15.14 classifies the automation track/point limits as runtime ceilings, contradicting 16.1's import-time semantic enforcement

Sections: 15.14 (line 1132), 16.1 (line 1318), §7 (line 285), 16.11 trace 6 (line 1483).

What the document says: 15.14: "The remaining PRD 58 limits — approximate one-shot voices (64), approximate continuous sounds (16), automation tracks (64), and automation points (256) — are runtime ceilings rather than document properties and belong to Phase 3c with the lifecycle contract." 16.1: "At most 64 automation tracks and 256 total automation points per expanded sound (PRD 58). Exceeding either is ERR_NODE_LIMIT_EXCEEDED" — a Semantic-stage code — and 16.11 trace 6 requires 65 tracks / 257 points to be rejected (at import, like the other trace-6 checks).

Why it is wrong: Track and point counts are fully decidable at import after deterministic component expansion; unlike voice ceilings they depend on nothing device- or runtime-specific. 15.14's "runtime ceilings rather than document properties" therefore contradicts 16.1, §7's Semantic filing, and trace 6, and leaves the two normative statements irreconcilable for an implementer deciding where a 65-track document fails. (The same stale sentence also mislabels only the track/point limits — the voice ceilings genuinely are runtime ceilings.)

Smallest fix: In 15.14, delete "automation tracks (64), and automation points (256)" from the runtime-ceilings sentence and add the two rows to the 15.14 authoring-time limits table (or add a pointer saying they are enforced as ERR_NODE_LIMIT_EXCEEDED per 16.1).


F6 — Automation point value sampling has no documented time, stream, or key

Sections: 16.1 (lines 1289, 1297), 14.4 (line 629), 14.6 (lines 652–658), 9.3 (line 457).

What the document says: Each point is { "at": <duration literal>, "value": ValueSpec<number> }, and "A track's values remain full ValueSpecs and may be random; it is only the curve's shape in time that is authored, not sampled."

Why it is wrong: random/choose ValueSpecs consume a seeded stream (9.3), and the document carefully fixes the sampling boundary for every other audio consumer: node fields resolve once at the sound instance's instantiation boundary in depth-first, property-document order (14.4), route depth resolves at the same boundary in route order (15.12), and sample-hold draws are documented with a child key on its own stream (14.12, and 9.3's "documented child key" requirement). Automation point values are none of these: a track lives on the graph object, its point values are not node fields, and 16.1 says nothing about when they are sampled or from which stream/key, nor where the automation array sits in the 14.4 sampling order. An implementer cannot reproduce a run that uses random automation point values, and 16.8's promise that a partially elapsed continuous sound keeps "a procedural sequence identical to an unbroken run" has no defined meaning for automated tracks. This is exactly the "value consuming a seeded random stream without a documented derivation key" class.

Smallest fix: One sentence in 16.1 stating that point value ValueSpecs resolve once at the owning sound instance's instantiation boundary, from that instance's own stream, sampled in depth-first, property-document order together with the graph's node fields and route depths (or on a documented per-track child key like 14.12's).


Sections: 16.5 (line 1387), 15.10 (lines 1020, 1024), 14.5 (lines 645–646).

What the document says: 16.5: a resonator contributes "the longest decay among its retained modes." 15.10: mode frequency may be authored up to audioMaxFrequency, and "A mode whose resolved frequency exceeds audioMaxFrequency at instantiation is omitted" with no diagnostic. 14.5: semantic checks use the device-independent ceiling 24000; the device ceiling is min(24000, sampleRate x 0.45).

Why it is wrong: On a 44.1 kHz device audioMaxFrequency is 19845 Hz; on 48 kHz it is 21600 Hz. A mode frequency in that gap (e.g. 20000 Hz or a ratio x fundamental product above the device ceiling) is import-legal but omitted at instantiation. A single-mode resonator then has zero retained modes, and "the longest decay among its retained modes" is undefined — so the one-shot ending bound, which must be "computable … from the resolved graph alone," has no defined value for a legal document on a legal device. (The node's output in that case is silence; only the bound row is left undefined.)

Smallest fix: Extend the row: "the longest decay among its retained modes, or zero when no mode is retained."


F8 — 14.9's impulse-envelope table contradicts its own formulas and its own prose

Section: 14.9 (lines 741–749).

What the document says: The table lists, for each decay mode, the envelope and a column "Value at p = 1":

decay Envelope Value at p = 1
flat a 0
linear a x (1 - p) 0
exponential a x e^(-6.907755 x p) (-60 dB at p = 1) 0

The prose immediately below reads: "Because flat and exponential do not reach zero on their own, the runtime applies a terminal linear fade to zero over the final min(1ms, d x 0.1) of the burst."

Why it is wrong: For flat the envelope is the constant a, so its value at p = 1 is a, not 0. For exponential, a x e^(-6.907755) is a x 0.001 — as the row's own "-60 dB at p = 1" annotation confirms — not 0. Only linear reaches 0 at p = 1. The column contradicts both the envelope formulas in the same table and the sentence that follows it (which exists precisely because those two envelopes do not reach zero). Note also that p = 1 is outside the stated domain 0 <= p < 1 of the envelope, so the column conflates the raw envelope with the post-fade value.

Smallest fix: Correct the two cells (e.g. flat → a; exponential → a x e^(-6.907755)), or delete the column entirely and keep the fade sentence as the definition at the burst end.


Defect classes with no findings

  • Prose contradicting a JSON example in the same section: none found. Every JSON example in 14–16 was checked against its field table and the section's rules and agrees with them (the 14.9 issue above is a table/formula/prose self-contradiction, not a JSON-example conflict).
  • Field introduced in one section but missing from its container's allowed-field table: none found. automation is in the 14.3 graph-object table; mode/release are in the 15.16 recipe table; parameters/input are in the 15.15 component table; the automation track fields (target, mode, interpolation, points) are all in the 16.1 table.
  • Cross-references to sections that do not exist: none found. All section references used in 14–16 (1.3, 4, 6.1/6.2, 7, 8.1/8.2, 9.3, 14.5, 15.13–15.17, 16.4, 16.8, and the "Phase 3c" deferrals) resolve to existing sections.
  • Diagnostic code used but absent from the section 7 table: none found. Every code used in 14–16 appears in §7, and the 15.18 and 16.10 tables duplicate §7 consistently. Wrong-stage filing is covered by F4 and F5.

Findings deliberately not reported

  • The provisional master-protection peak ceiling, tolerance, and release behavior (16.7), the GC4 audio long-stall bound (16.9), noise/impulse non-reproducibility (14.6), sample-hold's child-key documentation (14.12), and the audio.master absence (15.17) are all flagged in the text as deliberate and are consistent with the stated design decisions.
  • The ERR_TYPE_MISMATCH codes on out-of-set enum literals (e.g. 14.8 color: "grey", 15.3 mode: "comb") are internally consistent across the audio sections and defensible under section 2's definition of enum ("string constrained to an explicitly declared set of allowed tokens"), so no contradiction with §7 was established.