A Production-Oriented Framework for Evaluation of SFX Generation

Mélodie Desbos*, Yara Bahram, Eric Granger, Mohammadhadi Shateri
Dept. of Systems Engineering, LIVIA — École de Technologie Supérieure, Montréal, Canada
DAFx26 29th Int. Conf. on Digital Audio Effects · Cambridge, MA, USA
TL;DR — We propose a production-oriented evaluation framework for reference-guided SFX variation. It defines nine production requirements, evaluates five heterogeneous baselines under a shared ESC-50 audio-to-audio (ATA) protocol, and complements this with capability-specific analyses (morphing, temporal/energy alignment, inpainting, targeted editing). Among full-generation baselines, AudioX offers the strongest overall trade-off between reference alignment and diversity.
AudioX AudioLDM ThinkSound T-Foley A²SB
R1 Fidelity
R2 Identity
R3 Diversity
R4 Temporal
R5 Energy
R6 Control
R7 Targeted
R8 Robustness
R9 Efficiency
native/strong supported/moderate weak/not supported

Table 5 — Capability profile across methods.

Overview

In SFX production, sound designers often rely on a few reusable recordings, since manually designing variations is costly and time-consuming. This low variability causes perceptible repetition in interactive media, motivating systems that generate useful variations from a reference clip.

Existing methods are typically evaluated only within their original task settings (TTA, VTA, foley alignment, localized restoration), so their reported performance gives limited insight into controlled, production-useful, reference-guided variation — and makes them hard to compare.

Our framework = Requirements + Shared Protocol + Capability Profiles

  • 9 production requirements with operational definitions and failure modes.
  • A shared reference-guided ATA task on ESC-50 for fair comparison.
  • Capability-specific analyses that preserve each method's native strengths.

Production-oriented evaluation framework

Framework overview

A. A reference sound requires multiple diverse variations. B. Five baselines address complementary capabilities. C. A common objective with preserved capabilities, assessed via objective, perceptual, and signal-level diagnostics.

Nine Production Requirements

Each requirement has an operational definition, evaluation signals, and typical failure modes (full mapping in the capability profile below).

R1Fidelity & realism

FAD↓, MOS/S-MOS↑, spectrogram inspection. Fails via transient smearing, noise, synthetic texture.

R2Identity preservation

S-MOS↑, ImageBind alignment↑. Fails via semantic drift, generic textures, unrelated events.

R3Diversity without drift

ImageBind diversity↑, pairwise variation. Fails via near-duplicates or identity drift.

R4Temporal alignment

Onset error, FWHM ratio, energy-curve comparison. Fails via shifted onsets, wrong ordering.

R5Energy control

Envelope & energy-curve comparison, pre-onset Δ. Fails via loudness drift, distorted dynamics.

R6Controllability

Comparison across control settings (e.g. noise σ). Fails via weak response, over-transformation.

R7Targeted modification

Local qualitative + masked-region metrics. Fails via global rewriting, boundary discontinuities.

R8Robustness & stability

Cross-condition / mask-size / domain-shift analysis. Fails via out-of-domain breakdown.

R9Efficiency

Inference time, sampling steps, compute cost. Fails via excessive latency, impractical cost.

Five Heterogeneous Baselines

Selected to cover four distinct editing scopes: full-sample generation, temporally constrained generation, semantic reference-guided editing, and localized waveform inpainting.

AudioX

DiT · ICLR 2026

Multimodal (text/video/audio) anything-to-audio generation. Strong identity preservation, diversity without drift, and controllable reference-guided generation; supports SFX morphing via noise σ.

AudioLDM

Latent Diffusion · ICML 2023

VAE latent diffusion with CLAP conditioning. General-purpose full-sample variation, style transfer, and flexible text/reference control (R1, R3, R6).

ThinkSound

MMDiT (flow) · NeurIPS 2025

Semantic instructions + reference conditioning for targeted, object-centric editing while keeping plausible variation (R6, R7, R3).

T-Foley

Waveform Diff. + FiLM · ICASSP 2024

Sound-class and RMS-envelope conditioning for temporal-event control. Native temporal alignment and energy control (R4, R5).

A²SB

Attn. UNet (Schrödinger Bridge) · arXiv 2025 (NVIDIA)

Masks a segment and reconstructs it from surrounding context. Strictly local edits — relevant for targeted modification (R7) and context preservation (R2).

Two-stage protocol

(i) ATA generation — every method generates N=10 variants per reference under a shared ESC-50 setup (400 references → 4,000 outputs/model).
(ii) Method-specific analysis — each method is probed in its native editing setting.

Results

Capability profile

Compact summary of how each method fulfils the production requirements. Color encodes suitability:  native / strong  ·  supported / moderate  ·  weak / not supported .

Requirement AudioXAudioLDM ThinkSoundT-FoleyA²SB
R1 Fidelity & realism Strong; FAD 9.34, best S-MOS 3.37 Limited; FAD 20.09 Moderate; FAD 16.51 (8.97–9.30 local) Weak; FAD 24.53, noisy spectra Strong local; FAD 4.43 (inpaint)
R2 Identity preservation Strong; align 0.59, S-MOS 3.37 Limited; align 0.39, S-MOS 2.22 Moderate; align 0.51 (0.64–0.67 local) Weak; align 0.22, S-MOS 1.89 Strong local; S-MOS 4.81
R3 Diversity w/o drift Balanced; div 0.27, align 0.59 High but drift; div 0.43, align 0.39 Moderate; div 0.32, align 0.51 Drift-prone; div 0.32, align 0.22 Localized; div 0.28, mask-sensitive
R4 Temporal alignment Strongest; onset 43.52 ms Limited; onset 60.72 ms Moderate; onset 53.01 ms Mixed; explicit cond. but onset 63.53 ms Moderate local; onset 54.79 ms
R5 Energy control Implicit; no explicit control Limited; no energy condition Prompt-based; light local fluctuations Native; RMS-envelope conditioning Context-preserving; not explicit
R6 Controllability Strong; σ morphing 0.77→0.64 Strong but unstable; 0.40→0.15 Local but subtle edits Limited; class/RMS only Native mask control; 4.74→7.57
R7 Targeted modification Limited; full variation / morphing Limited; global style only Native; object-centric edits Not supported Native; masked inpainting
R8 Robustness & stability Good; 4.57→6.64 Sensitive; align 0.40→0.15 Stable but subtle Domain-sensitive; 13.51→25.75 Mask-sensitive; div 0.39→0.11
R9 Efficiency Moderate; ~400 variants/h Fast; ~5333 variants/h Fastest; ~6061 variants/h Slow; ~139 variants/h Slowest; ~121 variants/h

Local editing and inpainting results are diagnostic and not directly comparable to full-generation settings.

Shared ATA generation & transient diagnostics

Method Overall Quantitative Transient-Level Diagnostics
FAD↓S-MOS↑Div.↑Align.↑ FWHM →1Pre-onset Δ→0Onset Err. ms↓
AudioX (ICLR 2026) 9.343.37 [3.08, 3.66]0.270.59 0.985-0.00443.52
AudioLDM (ICML 2023) 20.092.22 [1.97, 2.48]0.430.39 0.969-0.00560.72
ThinkSound (NeurIPS 2025) 16.512.57 [2.25, 2.89]0.320.51 0.978-0.01153.01
T-Foley (ICASSP 2024) 24.531.89 [1.62, 2.16]0.320.22 1.0110.00863.53
Inpainting method — reported on inpainted regions only (not full-clip generation)
A²SB* (arXiv 2025) 4.434.81 [4.75, 4.87]0.280.49 0.9920.00554.79

Best and second best among full-generation methods. *A²SB metrics are computed on inpainted regions; all other metrics use cropped 4 s clips.

Diversity–identity trade-off

Diversity vs alignment

Higher diversity alone is insufficient when accompanied by weaker identity preservation. AudioX reaches the strongest alignment; AudioLDM the highest diversity.

Pareto-inspired analysis

Pareto analysis

Identity vs. temporal alignment (a) and identity vs. efficiency (b). Methods occupy distinct production trade-offs; higher is better on every axis.

Capability-specific diagnostics (native settings)

Setting Overall Quantitative Transient-Level Diagnostics
FAD↓S-MOS↑Div.↑Align.↑ FWHM →1Pre-onset Δ→0Onset Err. ms↓
A²SB inpainting (3 classes)
Mask 0.3–1.0 s4.740.390.350.97-0.0255.07
Mask 0.3–2.0 s7.570.110.310.890.0151.99
SFX morphing (3 classes) — ↓σ / ↑σ
AudioX4.57 / 6.640.10 / 0.260.77 / 0.640.95 / 0.870.007 / -0.0258.61 / 67.78
AudioLDM18.97 / 24.090.34 / 0.560.40 / 0.150.95 / 1.03-0.02 / -0.0364.82 / 86.78
T-Foley restricted pretraining manifold
Seen (5 classes)13.513.05 ± 0.780.340.340.970.0265.28
Unseen (45 classes)25.751.76 ± 0.620.320.211.010.00763.33
ThinkSound object-centric region editing (5 classes), mask 1.0–4.0 s
Attenuation9.060.190.671.130.01339.0
Enhancement9.300.210.641.150.0138.3
Reverberation8.970.220.641.120.0143.8

Diagnostic only — not intended for direct cross-method comparison.

Inference cost (R9)

MethodBackboneHardwareStepsBatchInferenceVariants/h
ThinkSoundMMDiTA100 40GB2420.66 h~6061
T-FoleyWave. Diff. + FiLMRTX A60002501628.8 h~139
A²SBAttn. UNetA100 40GB300133 h~121
AudioXDiTA100 40GB250110 h400
AudioLDMUNet (LDM)RTX A600020010.75 h~5333

Diagnostic indicator — methods use different backbones, batch sizes, and sampling steps, so values are not directly comparable.

Audio Demos

Reference-guided ATA variation

Given a reference clip and its class label, each method generates a variant.

SFX morphing

Reference-to-target transformation; AudioX preserves identity, AudioLDM transforms more strongly.

Localized inpainting (A²SB)

A masked region is reconstructed from the surrounding context.

Targeted editing (ThinkSound)

Object-centric edits over a masked [1 s; 3 s] region via semantic instructions.

Further Details

A. Objective evaluation & metrics

Distributional quality (R1) is measured with Fréchet Audio Distance (FAD) using the AudioLDM-eval implementation, computed per class and pooled. Because the shared ATA protocol is reference-conditioned rather than text-driven, identity (R2) and diversity (R3) use ImageBind audio embeddings: an audio-text metric such as CLAP would mainly reflect agreement with the class label rather than preservation of the specific reference event.

Alignment. With $\hat{z}^{ref}_r$ and $\hat{z}^{gen}_{r,v}$ the $\ell_2$-normalised embeddings of reference $r$ and variant $v$, alignment is the cosine similarity:

$$ A_{r,v} = \mathrm{cos\_sim}\!\left(\hat{z}^{ref}_r,\hat{z}^{gen}_{r,v}\right) = \frac{\langle \hat{z}^{ref}_r, \hat{z}^{gen}_{r,v}\rangle} {\lVert \hat{z}^{ref}_r\rVert_2\,\lVert \hat{z}^{gen}_{r,v}\rVert_2}. $$

Diversity. Following Seeing and Hearing, diversity is the mean pairwise semantic distance among the $V_r=10$ variants of a reference:

$$ D_r = \frac{2}{V_r(V_r-1)} \sum_{1\le i\lt j\le V_r} \left(1 - \mathrm{cos\_sim}\!\left(\hat{z}^{gen}_{r,i},\hat{z}^{gen}_{r,j}\right)\right). $$

High diversity must be read jointly with FAD and alignment, since a large spread can also indicate semantic drift. ImageBind alignment measures reference–variant similarity, while ImageBind diversity measures variation among generations from the same reference; temporal behaviour is evaluated separately via signal-level onset diagnostics below.

B. Transient-level diagnosis

From the log-mel spectrogram $\ell_{m,t}$ we keep only positive temporal increases, sum over mel bins, and normalise to $[0,1]$ to obtain an onset-strength envelope $\tilde{e}_t$ focused on transient shape rather than loudness:

$$ D_{m,t} = \max(\ell_{m,t+1}-\ell_{m,t},\,0), \qquad e_t = \sum_m D_{m,t}, \qquad \tilde{e}_t = \frac{e_t - \min_t e_t}{\max_t(e_t-\min_t e_t)+\epsilon}. $$

Peaks are detected with a 40 ms minimum spacing and prominence $\rho=0.10$, yielding three diagnostics.

FWHM ratio — transient width preservation (→1 ideal; >1 smeared, <1 sharpened):

$$ \mathrm{FWHM}(p)=\omega_p\,\Delta_{\mathrm{ms}}, \qquad R_{\mathrm{FWHM}}=\frac{\mathrm{FWHM}_{\mathrm{gen}}}{\mathrm{FWHM}_{\mathrm{ref}}+\epsilon}. $$

Pre-onset Δ — energy before vs. after the transient (40 ms vs. 80 ms windows; →0 ideal, positive = pre-echo):

$$ R_{\mathrm{pre}}(p)=\frac{\sum_{t=\max(0,p-q_{\mathrm{pre}})}^{p-1}\tilde{e}_t} {\sum_{t=p}^{\min(T,p+q_{\mathrm{post}})-1}\tilde{e}_t+\epsilon}, \qquad \Delta_{\mathrm{pre}}=R_{\mathrm{pre,gen}}-R_{\mathrm{pre,ref}}. $$

Onset error — median timing gap between matched reference and generated onsets (tolerance $\delta_{\max}=150$ ms):

$$ j^*(i)=\arg\min_{j\in U}\bigl|\tau^{gen}_j-\tau^{ref}_i\bigr|, \qquad E_{\mathrm{onset}}=1000\cdot\mathrm{median}_i\bigl|\tau^{gen}_{j^*(i)}-\tau^{ref}_i\bigr|. $$

C. Subjective evaluation (listening study)

A Similarity Mean Opinion Score (S-MOS) was rated by 15 participants on a 1–5 identity scale over reference–variation pairs drawn from the 4,000 outputs per method. Participants used headphones in a quiet environment, could replay each pair freely, and rated 100 randomised trials per method. Trials were anonymised and randomised; scores are reported with 95% confidence intervals. The prompt was: "For each trial: listen to the Reference, then the Candidate. Rate the identity fidelity (1–5), defined as the similarity to the reference event/source (excluding loudness)."

D. Additional training details

All baselines are initialised from released pretrained checkpoints and adapted to ESC-50 with deliberately lightweight fine-tuning, keeping encoders frozen when applicable and adapting only task-relevant generation/conditioning layers. The ESC-50 test fold has 50 classes × 8 references; for each reference, $N=10$ variants are generated (4,000 outputs/model).

  • Latent diffusion (AudioLDM, ThinkSound): text/audio encoders frozen; only the diffusion backbone (UNet/MMDiT) and conditioning projection layers fine-tuned. ThinkSound: 10 epochs × 150 steps; AudioLDM: 200 steps. Updates kept low to limit overfitting and drift.
  • Waveform diffusion (T-Foley): reduced from 500 to 25 epochs × 250 steps (6,250 steps); class/MLP embeddings and FiLM layers updated; operates at 22 kHz (native rate preserved).
  • AudioX: 20 epochs × 200 steps (4,000 steps) via stable-audio-tools. Fixed 11 s window — each 5 s clip zero-padded with a padding-mask loss; conditioning uses only the class name (audio/video modalities empty).
  • A²SB: two runs matched to pretrained masking splits (0.0–0.5 s, 0.5–1.0 s) plus extended masked-duration fine-tuning (short 0.3–1.0 s, long 0.3–2.0 s). STFT features with n_fft=2048, hop=512, 44.1 kHz.
  • Sampling rates: native rates preserved during fine-tuning (ThinkSound 44 kHz, T-Foley 22 kHz, A²SB 44.1 kHz); resampling to 16 kHz only at evaluation when needed for fair comparison.

BibTeX

@inproceedings{desbos2026sfxeval,
  title     = {A Production-Oriented Framework for Evaluation of SFX Generation},
  author    = {Desbos, M\'elodie and Bahram, Yara and Granger, Eric and Shateri, Mohammadhadi},
  booktitle = {Proc. 29th Int. Conf. on Digital Audio Effects (DAFx)},
  address   = {Cambridge, MA, USA},
  year      = {2026}
}