Configuration Reference
Interactive builder and full reference for every option in the Python SDK.
Every parameter maps directly to client.process().
Cleanvoice resamples all audio to 44.1 kHz or 48 kHz before processing. Your original file is not modified — the cleaned output uses the resampled rate.
Option reference
Audio cleaning
| Option | Type | Default | Description |
|---|---|---|---|
fillers | bool | False | Remove "um", "uh", "like", and similar filler words |
long_silences | bool | False | Trim long pauses and gaps between sentences |
mouth_sounds | bool | False | Remove clicks, lip smacks, and tongue sounds |
breath | bool | str | False | Remove audible breathing between sentences |
stutters | bool | False | Remove repeated word fragments ("I— I— I think") |
hesitations | bool | False | Remove short hesitation sounds that aren't full filler words |
muted | bool | False | Silence edits instead of cutting — preserves original timing |
Filler and hesitation detection is language-aware. English, German, and Romanian have the most accurate models. See supported languages.
breath options:
| Value | Behavior |
|---|---|
True | Recommended for most audio. Best default for challenging recordings. |
"legacy" | Conservative removal. Safer choice for already-clean recordings. |
"natural" | Lighter touch — preserves more of the original breathing feel. |
False | Disabled (default). |
Audio enhancement
Recommended default: remove_noise=True and normalize=True. True, and omitting remove_noise, runs Noise v2. "legacy" pins the previous engine (DeepFilterNet). "v2" is still accepted and means the same as True. normalize=True is Normalize v2 (a boolean; there is no "v2" string). Filler, silence, and breath options stay off unless you ask for them.
result = client.process("episode.mp3", remove_noise=True, normalize=True)| Option | Type | Default | Description |
|---|---|---|---|
remove_noise | bool | str | True (Noise v2) | True is Noise v2 (recommended). Omitting it is the same as True. "legacy" pins the previous engine (DeepFilterNet). "v2" is still accepted. False turns it off. |
studio_sound | bool | str | False | Optional, more aggressive enhancement. |
normalize | bool | False | Normalize v2 when True. Levels loudness. Off when omitted. |
keep_music | bool | False | Preserve music sections during noise reduction |
autoeq | bool | False | Legacy automatic EQ option. Prefer studio_sound; autoeq will be removed in a future release. |
remove_noise values:
| Value | Behavior |
|---|---|
True | Recommended. Noise v2. This is also the default when the argument is omitted. |
"legacy" | Previous engine (DeepFilterNet), pinned so it stays on that model. |
"v2" | Still accepted. Same as True (Noise v2). |
False | Off. |
studio_sound options:
| Value | Behavior |
|---|---|
True | Aggressive studio-quality enhancement. Optional, and stronger than Remove Noise v2. |
"nightly" | Advanced/experimental variant. Currently behaves similarly to True. |
False | Disabled (default). |
Output
| Option | Type | Default | Description |
|---|---|---|---|
export_format | str | "auto" | Audio-only output format: mp3, wav, flac, m4a, or auto (matches input). Video jobs keep the original video container format. |
target_lufs | float | omitted (-16 with Normalize v2) | Integrated loudness when normalize=True. Accepted range is -36 to -6. -16 is the podcast target and the value Normalize v2 uses when this argument is omitted. |
mute_lufs | float | -120 | Gate level for LUFS measurement. Default -120 disables gating. |
export_timestamps | bool | False | Return a JSON file with edit markers for use in a DAW or NLE |
To set a loudness target, pass target_lufs to the value you want:
result = client.process("episode.mp3", target_lufs=-16.0)Content generation
| Option | Type | Default | Description |
|---|---|---|---|
transcription | bool | False | Full word-by-word transcript. Language auto-detected. |
summarize | bool | False | Chapter markers, key learnings, and episode summary. Enables transcription automatically. |
social_content | bool | False | Generate tweets, LinkedIn posts, and show notes. Enables summarize automatically. |
Advanced
| Option | Type | Default | Description |
|---|---|---|---|
signed_url | str | None | Pre-signed PUT URL. Cleanvoice uploads directly to your storage instead of hosting the file. |
merge | bool | False | Multi-track only. Merge all tracks into a single output file. |
audio_for_edl | bool | False | Video workflows only. Return additional uncut enhanced audio alongside the edited video for EDL/NLE work. |
video | bool | False | Process the input as video. The SDK auto-detects common video filenames and URLs, but explicit video=True is safest for ambiguous or extensionless URLs. |
For raw REST requests, video must be set to true for actual video editing. Otherwise the file is treated as audio-only. In the Python SDK, common video paths are auto-detected, but being explicit is still safest.
SDK-only options
| Option | Type | Default | Description |
|---|---|---|---|
output_path | str | None | Save the cleaned audio to this path automatically |
progress_callback | callable | None | Called with a dict payload during polling: status, result, edit_id, and attempt |
def on_progress(update):
print(f"Status: {update['status']}")
result = client.process(
"episode.mp3",
remove_noise=True,
normalize=True,
progress_callback=on_progress,
output_path="cleaned.mp3",
)