Audio as Code

Reference

Audio as Code 0.2.0. Python distribution audio-as-code, import audio_as_code, command aac. Install audio-as-code==0.2.0 from PyPI for the engine and CLI; use the source workspace for the complete examples and portable skill. See the quickstart for both routes. The 0.2 API may change; breaking score format changes will use a new schema_version. Expression fields are optional additions to version 1.

Machine-readable versions of this page: the score schema at ../schemas/song-v1.schema.json (same as aac schema) and the instrument catalog at ../instruments.json (same as aac instruments --all).

Command line#

text
aac [-h] [--version] {init,doctor,instruments,demo,schema,validate,inspect,render,preview,midi,analyze} ...

Also runnable as python -m audio_as_code. With uv, prefix with uv run.

CommandArgumentsDoes
aac --versionPrint {"version": "0.2.0"}
aac initDIRECTORYCreate a starter in a new or empty folder; never install dependencies
aac doctorCheck the imported runtime and short synthesis path
aac instruments[--all] [--family F] [--engine E]Print the catalog. Without --all, only playable voices
aac demo[-o OUTPUT] (default demo.json)Write the built-in demo score
aac schema[-o OUTPUT]Print the JSON Schema, or write it to a file
aac validatescoreCheck a JSON score; print a summary
aac inspectscoreInspect tracks and WAV/MIDI readiness without rendering or writing files
aac render`score -o OUTPUT [--stems DIRECTORY] [--format pcm16\pcm24\float32] [--report REPORT]`Render stereo WAV; PCM16 by default
aac previewscore -o OUTPUT --start SECONDS --duration SECONDSExport an excerpt after a full-context render
aac midiscore -o OUTPUTExport a type-1 MIDI file
aac analyzeaudioMeasure 16/24-bit PCM or 32-bit float WAV

--family choices: plucked_strings, bowed_strings, keyboards, woodwinds, brass, percussion, pitched_percussion, electronic. --engine choices: string, air_column, modal, membrane, electronic.

Exit status and output. Success: exit 0, one JSON object on stdout. Failure: exit 2, one JSON object on stderr, either {"error": "invalid_score", "issues": [{"path", "message", "type"}]} or {"error": "operation_failed", "message": "..."}, optionally with a hint string. -h/--help prints plain text, not JSON.

Render options. Both render commands accept --format, --report, --progress-file PATH (live JSON Lines), and either --no-normalize or --target-lufs FLOAT. A target requires audio-as-code[loudness]; --peak-ceiling-dbfs FLOAT requires a target and defaults to -1 dBFS. It is a sample-peak bound, not a true-peak limiter. Preview costs a full render. Read production output for exact semantics.

Files. Parent directories are created. Existing outputs are overwritten. The score, WAV, report, progress log, and every stem path (DIRECTORY/01.wav, 02.wav, … in track order) must be distinct.

Inspection. inspect exits 0 when inspection succeeds, including when the score cannot be exported. Check readiness.render.ready and readiness.midi.ready. Each issue has code, severity, target, path and message. Only issues with severity: "error" block their export target. Readiness covers static score rules; it does not test output paths, runtime resources or audio quality. Polyphony counts written note intervals, including muted notes, and excludes releases and effects.

Score format (schema version 1)#

Unknown fields are rejected everywhere. Numbers must be finite. Times are quarter-note beats.

Song#

FieldTypeDefaultConstraint
schema_version"1""1"Only "1"
titlestring"Untitled"1–200 characters
bpmnumber12020–300
beatsnumber16> 0, ≤ 65536; every note must end by it
sample_rateinteger4410022050, 44100, or 48000
seedinteger00–4294967295
master_gainnumber0.80–1
tracksarray of Trackrequired1–64, unique names, ≤ 100,000 notes in total
tempo_maparray of TempoChange[]At most 1024; strictly increasing beats before song end
automationarray of Automation[]At most one lane, targeting master_gain
effectsarray of Delay or Reverb[]Ordered chain, at most 4 effects

Track#

FieldTypeDefaultConstraint
namestringrequired1–80 characters, at least one visible character, unique
instrumentstring"pluck"A playable catalog ID
gainnumber0.60–1, linear
pannumber0−1 (left) to 1 (right), equal-power
toneTone or nullnullOnly fields in the instrument's tone_controls
articulation"soft", "accented", or nullnullOnly values in the instrument's articulations; null preserves original sound
notesarray of Note[]drum_machine pitches must be 36, 38, 42, 45, 49, 54, 60, or 64
release_secondsnumber00–10; inherited by notes without an override
pedalarray of PedalEvent[]Piano only when nonempty; at most 1,024 ordered alternating down/up events within the song
automationarray of Automation[]At most two lanes; unique gain/pan targets
effectsarray of Delay or Reverb[]Ordered chain, at most 4 effects

Track.pedal defaults to [] and supports up to 1,024 ordered PedalEvent objects on piano. Each has a nonnegative beat and strict boolean down. Events alternate down/up, start down and cannot exceed song end. Notes released under the pedal ring until lift; an open pedal lifts at score end. A default 0.12-second damper release applies to captured notes unless overridden. See the pedal timing, release and MIDI rules.

Note#

FieldTypeDefaultConstraint
pitchinteger or string"C4"MIDI 0–127, or a name: letter A–G (either case), optional # or b, octave −1 to 9 (C4 = 60, C-1 = 0)
startnumber0≥ 0
durationnumber1> 0
velocitynumber0.8> 0, ≤ 1
release_secondsnumber or nullnull0–10; null inherits the track; zero keeps the original gate

| articulation | "soft", "accented", or null | null | Null inherits the track; a value overrides it |

Read articulations for the 13 supported bowed-string and wind voices, note inheritance, designed attack/release gestures and MIDI limitations.

Tempo, automation, and effects#

TempoChange(beat, bpm) uses nonnegative beats and BPM 20–300. Changes are steps; bpm applies before the first change, and a change at beat zero overrides it. Note starts and ends are converted by integrating the tempo segments. Continuous tempo ramps are not implemented.

Automation(parameter, points, interpolation="linear") takes 1–1024 AutomationPoint(beat, value) entries with strictly increasing, nonnegative beats at or before song end. interpolation is "linear" or "step". Values are absolute: gain/master_gain 0–1, pan -1–1. Before the first and after the last point, its endpoint value holds. Linear interpolation follows beats even across tempo changes. Song lanes only accept master_gain; track lanes accept gain and pan. Each parameter can appear only once per owner.

Effect fieldDefaultConstraint
Delay type"delay"JSON discriminator
Delay time_seconds0.250.01–2
Delay feedback0.40–0.85; decay between successive echoes
Delay repeats4Integer 1–16; finite echo train
Delay mix0.20–1 dry/wet mix
Reverb type"reverb"JSON discriminator
Reverb decay_seconds1.20.1–5
Reverb mix0.20–1 dry/wet mix

Effects use generated delay/comb/diffusion processing, not measured impulse responses. Track effects precede master gain; song effects process the summed mix. Effects and note releases reserve automatic, bounded tails. Releases use a cosine fade after note-off; zero preserves the old duration-gated sound. Note on/off positions must still fit within beats.

Tone#

FieldDefaultRange
brightness0.50–1
decay_secondsnull (catalog default)0.1–20
pluck_positionnull0.05–0.45
breathnull0–1
vibrato_depth_centsnull0–100
vibrato_rate_hznull0.1–12
glide_semitonesnull−24–24
detune_centsnull0–40

null selects the instrument's default. Which instruments accept which fields is in the composition guide and in each catalog entry's tone_controls.

Python API#

Everything below is importable from audio_as_code.

Project setup and arrangement#

init_project(path) and doctor() return JSON-compatible dictionaries; see project setup. The setup command uses packaged templates and does not need the repository.

beat_at_seconds(song, seconds) inverts step tempo changes. place_at_seconds(song, pattern, seconds) returns placed notes. Arrangement(song, sections) holds named Section(name, start_beat, end_beat) intervals; .replace_section(name, {track_name: pattern}) returns a revised arrangement with the same section length. LoopRegion(start_beat, end_beat) and render_loop_preview(song, region, repetitions=2) produce repeated audio and join measurements. See arrangement for boundary/tail rejection, note-index reseeding and the difference between repeated audio and continuing effect state. These helpers do not add fields to score JSON.

Score models#

Song, Track, Note, Tone, TempoChange, PedalEvent, Automation, AutomationPoint, Delay, and Reverb are frozen Pydantic models with the fields above. Construct them with keyword arguments or validate data:

CallResult
Song(...), Song.model_validate(dict)Validated song; raises pydantic.ValidationError
Song.load(path)Read and validate a JSON file
song.save(path)Write indented JSON, creating parent directories
song.secondsScore duration, integrating the tempo map; beats * 60 / bpm without one
song.beat_to_seconds(beat)Time at a nonnegative beat; last tempo continues beyond score end
song.render_secondsScore duration plus any required release/effect tail
song.model_dump(mode="json")Plain data for editing

midi_pitch(value) -> int converts a pitch name or integer to MIDI and raises ValueError if invalid.

Pattern#

Pattern(notes, beats) is an immutable dataclass. Build one with Pattern.sequence(steps, *, step=1, gate=0.8, velocity=0.8) where each step is a pitch, a list or tuple of pitches (a chord), or None (a rest). gate must be in (0, 1].

MethodReturns
.repeat(times: int)Pattern, times ≥ 1
.transpose(semitones: int)Pattern
.then(other: Pattern)Pattern
.overlay(other: Pattern, *, offset=0)Pattern; layer at a nonnegative beat offset, retaining all notes and full phrase lengths
.stretch(factor)Pattern; scale beat timing and length by a positive factor; releases stay in seconds
.scale_velocity(factor)Pattern; multiply velocities by a positive factor, with every result in (0, 1]
.at(beat: float)tuple[Note, ...] on the absolute timeline

Rendering#

FunctionReturns
render(song, path, *, normalize=True, stems_dir=None)Report dict; writes stereo WAV and optional stems (PCM16 by default)
render_audio(song, *, normalize=True)Object with .audio (float32 array, frames × 2) and .report; writes nothing
analyze_wav(path)Measurement dict for PCM16/24 or float32 WAV
inspect_score(song)Static score summary, per-track measurements, export readiness and structured issues; no audio or file writes
export_midi(song, path)MIDI summary dict

Renders are limited to 300 seconds including release/effect tails. Default normalization only attenuates the mix to a 0.95 sample peak. A loudness target may boost or attenuate instead. With normalize=False, values beyond full scale clip in PCM exports; float32 exports preserve finite overs. Stems contain track processing and master gain/automation, but omit master effects; a warning explains when their sum differs from the processed mix.

All render functions also accept progress=callback, cancel=predicate, target_lufs=None, and peak_ceiling_dbfs=-1.0. File exporters accept wav_format="pcm16". render_preview_audio and render_preview require start_seconds and duration_seconds and retain full-song processing context. RenderProgress carries stage counts; RenderCancelled signals a cooperative stop. measure_loudness(audio, sample_rate) needs the optional loudness extra. Production output specifies gain, metering and publication behavior; default rendering does not import the optional backend.

Catalog#

FunctionReturns
instrument_catalog(*, family=None, engine=None, include_planned=False)The aac instruments dict
list_instruments(*, family=None, engine=None, include_planned=False)Tuple of InstrumentInfo
get_instrument(id)One InstrumentInfo; ValueError if unknown

InstrumentInfo fields: id, name, family, engine, status (available or planned), description, tone_controls, default_decay_seconds, midi_program, midi_note, preview_pitch, default_tone, articulations.

Render report#

Returned by render() and printed by aac render.

FieldMeaning
engine_version, numpy_versionInstalled versions used
score_sha256Hash of the canonical validated score
title, tracks, notes, seed, normalizeFrom the score and options
gain_appliedMix attenuation, ≤ 1
score_duration_seconds, tail_secondsTempo-integrated score length and additional reserved audio tail
effectsMaster and per-track effect settings
before_gain, audioFloat measurements before and after gain
output, wavWAV path and measurements of the saved file
stems[{"track", "path", "audio"}], empty without stems
warningsHuman-readable strings

Float measurements: sample_rate, channels, frames, duration_seconds, peak, rms, peak_dbfs, rms_dbfs (null when silent), clipped_samples (channel samples with |x| ≥ 1), silent (exact digital zero). WAV measurements replace clipped_samples with full_scale_samples (PCM endpoint samples, or float samples with absolute amplitude at least 1). These are signal checks only.

MIDI export#

Piano pedal events export binary sustain CC64 values 127/0 while retaining written note-on/off times. Pedal changes precede note-offs at the same tick; an open pedal emits CC64=0 at score end. Damper sound depends on the receiving synthesizer.

Type 1, 480 ticks per quarter note, one tempo track plus one track per score track, with names, General MIDI programs (approximate mappings from the catalog), volume (CC 7 = gain × master gain), and pan (CC 10). Percussion IDs and drum_machine share channel 10, and their note events are written into the first percussion track (other percussion track names remain as metadata); drum gain is folded into velocity and drum pan is not exported. Limits: at most 15 melodic tracks, one track per individual drum ID, no note shorter than one tick, and no overlapping notes of the same pitch on a channel (including a kit and an individual drum). Timing can differ from the WAV by up to a tick. MIDI carries step tempo changes on the tick grid but does not carry the synthesizers, Tone settings, articulations, automation, effects, note releases, audio tails or normalization; how it sounds depends on the receiver.

Synthesis scope#

All 49 catalog instruments are generated from code across five engine groups: vibrating strings, air columns, bars/bells/plates, membranes, and electronic synthesis. There are no recorded samples, SoundFonts, or measured impulse responses. The models are designed approximations. Bowed strings and winds use harmonic source/filter models without bow-friction or reed/bore physics. Piano supports binary damper-pedal gates; half-pedaling, sympathetic resonance and note-to-note articulations are not modeled. Each catalog description states that instrument's scope.

Reproducibility#

Identical score, software versions, and platform produce identical WAV bytes. Each note's generated noise is seeded from the song seed, track name, and note index, so adding a track does not change the others. Keep the score, uv.lock (or your installed versions), and the render report together.

Repository files#

Path in the source folderContents
src/audio_as_code/The library (model.py, pattern.py, render.py, midi.py, instruments.py, cli.py, synthesis modules)
schemas/song-v1.schema.jsonChecked-in score schema
docs/site/These pages, llms.txt, and the tutorial examples
skills/audio-as-code/SKILL.mdPortable agent setup, composition and revision workflow
examples/Showcase, instrument browser, and other render scripts
LICENSEMIT license
CONTRIBUTING.mdChecks to run before contributing
Next: Runnable examples