DeepF0: End-to-end fundamental frequency estimation for music and speech signals

Satwinder Singh; Ruili Wang; Yuanhang Qiu

doi:10.1109/ICASSP39728.2021.9414050

DeepF0: End-to-end fundamental frequency estimation for music and speech signals

Satwinder Singh, Ruili Wang, Yuanhang Qiu

Research output: Journal Publication › Conference article › peer-review

26 Citations (Scopus)

Abstract

We propose a novel pitch estimation technique called DeepF0, which leverages the available annotated data to directly learns from the raw audio in a data-driven manner. f0 estimation is important in various speech processing and music information retrieval applications. Existing deep learning models for pitch estimations have relatively limited learning capabilities due to their shallow receptive field. The proposed model addresses this issue by extending the receptive field of a network by introducing the dilated convolutional blocks into the network. The dilation factor increases the network receptive field exponentially without increasing the parameters of the model exponentially. To make the training process more efficient and faster, DeepF0 is augmented with residual blocks with residual connections. Our empirical evaluation demonstrates that the proposed model outperforms the baselines in terms of raw pitch accuracy and raw chroma accuracy even using 77.4% fewer network parameters. We also show that our model can capture reasonably well pitch estimation even under the various levels of accompaniment noise.

Original language	English
Pages (from-to)	61-65
Number of pages	5
Journal	Proceedings - ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing
Volume	2021-June
DOIs	https://doi.org/10.1109/ICASSP39728.2021.9414050
Publication status	Published - 2021
Externally published	Yes
Event	2021 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2021 - Virtual, Toronto, Canada Duration: 6 Jun 2021 → 11 Jun 2021

Keywords

F0 estimation
Pitch estimation
Speech processing
Temporal convolutional network

ASJC Scopus subject areas

Software
Signal Processing
Electrical and Electronic Engineering

Access to Document

10.1109/ICASSP39728.2021.9414050

Cite this

@article{a0e5aedfd42e4efab5c14f1ccd5b5f0e,

title = "DeepF0: End-to-end fundamental frequency estimation for music and speech signals",

abstract = "We propose a novel pitch estimation technique called DeepF0, which leverages the available annotated data to directly learns from the raw audio in a data-driven manner. f0 estimation is important in various speech processing and music information retrieval applications. Existing deep learning models for pitch estimations have relatively limited learning capabilities due to their shallow receptive field. The proposed model addresses this issue by extending the receptive field of a network by introducing the dilated convolutional blocks into the network. The dilation factor increases the network receptive field exponentially without increasing the parameters of the model exponentially. To make the training process more efficient and faster, DeepF0 is augmented with residual blocks with residual connections. Our empirical evaluation demonstrates that the proposed model outperforms the baselines in terms of raw pitch accuracy and raw chroma accuracy even using 77.4% fewer network parameters. We also show that our model can capture reasonably well pitch estimation even under the various levels of accompaniment noise.",

keywords = "F0 estimation, Pitch estimation, Speech processing, Temporal convolutional network",

author = "Satwinder Singh and Ruili Wang and Yuanhang Qiu",

note = "Publisher Copyright: {\textcopyright} 2021 IEEE; 2021 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2021 ; Conference date: 06-06-2021 Through 11-06-2021",

year = "2021",

doi = "10.1109/ICASSP39728.2021.9414050",

language = "English",

volume = "2021-June",

pages = "61--65",

journal = "Proceedings - ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing",

issn = "1520-6149",

publisher = "Institute of Electrical and Electronics Engineers Inc.",

}

TY - JOUR

T1 - DeepF0

T2 - 2021 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2021

AU - Singh, Satwinder

AU - Wang, Ruili

AU - Qiu, Yuanhang

PY - 2021

Y1 - 2021

N2 - We propose a novel pitch estimation technique called DeepF0, which leverages the available annotated data to directly learns from the raw audio in a data-driven manner. f0 estimation is important in various speech processing and music information retrieval applications. Existing deep learning models for pitch estimations have relatively limited learning capabilities due to their shallow receptive field. The proposed model addresses this issue by extending the receptive field of a network by introducing the dilated convolutional blocks into the network. The dilation factor increases the network receptive field exponentially without increasing the parameters of the model exponentially. To make the training process more efficient and faster, DeepF0 is augmented with residual blocks with residual connections. Our empirical evaluation demonstrates that the proposed model outperforms the baselines in terms of raw pitch accuracy and raw chroma accuracy even using 77.4% fewer network parameters. We also show that our model can capture reasonably well pitch estimation even under the various levels of accompaniment noise.

AB - We propose a novel pitch estimation technique called DeepF0, which leverages the available annotated data to directly learns from the raw audio in a data-driven manner. f0 estimation is important in various speech processing and music information retrieval applications. Existing deep learning models for pitch estimations have relatively limited learning capabilities due to their shallow receptive field. The proposed model addresses this issue by extending the receptive field of a network by introducing the dilated convolutional blocks into the network. The dilation factor increases the network receptive field exponentially without increasing the parameters of the model exponentially. To make the training process more efficient and faster, DeepF0 is augmented with residual blocks with residual connections. Our empirical evaluation demonstrates that the proposed model outperforms the baselines in terms of raw pitch accuracy and raw chroma accuracy even using 77.4% fewer network parameters. We also show that our model can capture reasonably well pitch estimation even under the various levels of accompaniment noise.

KW - F0 estimation

KW - Pitch estimation

KW - Speech processing

KW - Temporal convolutional network

UR - http://www.scopus.com/inward/record.url?scp=85115175742&partnerID=8YFLogxK

U2 - 10.1109/ICASSP39728.2021.9414050

DO - 10.1109/ICASSP39728.2021.9414050

M3 - Conference article

AN - SCOPUS:85115175742

SN - 1520-6149

VL - 2021-June

SP - 61

EP - 65

JO - Proceedings - ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing

JF - Proceedings - ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing

Y2 - 6 June 2021 through 11 June 2021

ER -

DeepF0: End-to-end fundamental frequency estimation for music and speech signals

Abstract

Keywords

ASJC Scopus subject areas

Access to Document

Other files and links

Fingerprint

Cite this