Time-Unconditional Generative Speech Enhancement via Autonomous Rectified Flow

Wen Zhang, Wenbin Jiang, Yang Zhang, Xiaofei Zhou

Abstract

Most generative speech enhancement methods rely on explicit time-step embeddings for temporal conditioning. In this paper, we propose the Autonomous Rectified Flow framework, which challenges the necessity of such conditioning. Using a linear interpolation path, we show that the target vector field is inherently time-invariant. We further introduce a time-unconditional network that eliminates explicit time-step information and infers the denoising direction solely from the spatial relationship between the current state and the noisy observation. Predicting this target vector field is equivalent to modeling the noise distribution. By avoiding overfitting to temporal trajectories, the proposed autonomous design significantly improves generation quality, robustness, and inference efficiency.

Dataset

The proposed framework is evaluated on the VoiceBank+DEMAND dataset.

1. Audio Performance Comparison

Listen and compare the speech quality under different models and inference steps (NFE).

Sample ID Noisy Clean FLOWSE (Baseline) ARFSE (Proposed)
NFE = 1 NFE = 5 NFE = 1 NFE = 5
p232_250
p232_409
p257_008
p257_033
p257_106

2. Spectrogram Analysis (Case Study: p257_008)

Visual comparison of spectrogram recovery across FLOWSE baseline and our proposed ARFSE framework.

Spectrogram Comparison for p257_008