DeepPitch: Wide-Range Monophonic Pitch Estimation using Deep Convolutional Neural Networks

Abstract – Pitch Estimation is an important problem in Machine Hearing, with application in Music Information Retrieval, Speech Analysis, and Auditory Scene Analysis. In this paper, we propose DeepPitch, a Deep Convolutional Neural Network (DCNN) approach that operates on a high-resolution log-frequency spectrum (9 taps/semitone), with one DCNN operating as a binary Pitch Activity Detector, and another DCNN estimating the Pitch Value at 9 taps/semitone resolution. Most Pitch Estimation algorithms such as pYIN, CREPE, RAPT, etc., use a restricted range of output pitch values, compared to the wide range of human pitch perception. DeepPitch was designed to cover a wide pitch range of 27.5 Hz to 4.186 kHz, corresponding to the full range of a piano. The training and test data are based on a 107-minute set of 44.1 kHz audio samples including speech (male, female, child), musical instruments (piano, guitar, bass, violin), synthetic test signals (sine waves, missing fundamental examples, sirens), and bandlimited speech and piano. Performance of the algorithm is 99.75% accuracy on Pitch Activity Detection, 96.92% on Pitch Value.

1. Introduction

Pitch Estimation, or fundamental frequency (f0) estimation, of a monophonic audio signal is an important problem in Music Information Retrieval, Speech Analysis, and Auditory Scene Analysis, and monophonic Pitch Estimation is a necessary prerequisite for the development of polyphonic Pitch Estimation. Many front-end signal representations have been used in many different Pitch Estimation algorithms, including linear-frequency spectral frames, cepstrum, autocorrelation function (ACF), average magnitude difference function (AMDF), normalized cross-correlation function (NCCF), and cumulative mean normalized difference function (CMNDF).

In this paper, we present DeepPitch, a monophonic Pitch Estimation method using a high-resolution log-frequency spectrum as the front-end signal representation. The training data was chosen to adapt the system to a broad set of stimuli, covering the full 27.5 Hz to 4.186 kHz range of a piano: speech (male, female, child), musical instruments (piano, bass, guitar, violin), synthetic test signals (sine waves, missing fundamental examples, sirens), and bandlimited speech and piano.

2. Architecture

The DeepPitch architecture is shown in Figure 1.

The Pitch Activity Detector DCNN consists of two convolutional layers and two fully connected layers. The first convolutional layer uses 16 1-dimensional kernels of size 108 frequency values. This convolution operates on the 1044 frequency values and produces 16x1044 outputs, which are then rectified and max-pooled to 16x522. The second convolutional layer uses 4 2-dimensional kernels of size 16x108. This convolution produces 4x522 outputs, which are then rectified and max-pooled to 4x261, and then flattened to 1044 outputs. The third layer is fully connected with 1044 inputs and 256 rectified outputs, followed by a dropout layer with dropout probability 0.5. The fourth layer is fully connected with 256 inputs and 2 outputs, with a SoftMax, representing a binary Pitch Activity Detection.

The Pitch Value Estimator DCNN consists of two convolutional layers and two fully connected layers. The architecture is similar, with the main difference being the number of output nodes in the final layer, which is 793 outputs, covering the 88-semitone range of a piano at 9 taps/semitone.

Both networks are trained to minimize the binary cross-entropy between the target vector and the predicted vector. This loss function is optimized using the ADAM optimizer, with a learning rate of 0.0001.

3. Experiments

3.1. Data Sets

The Core Data Set is 8 minutes and 40.04 seconds of audio, containing speech, full scales played on musical instruments, sine wave scale, and test signals. The ground truth targets were produced using get_f0 (RAPT) from WaveSurfer software, and subsequently corrected. An Extended Data Set was created by pitch-shifting the speech and musical instrument log-frequency spectral values through a range of ±4 frequency taps to fill in the unrepresented frequency taps.

3.2. Methodology

We trained the model using a simple 70/10/20 train/validation/test split. For evaluation of the performance of the pitch estimator, an estimate was considered correct if it matched the target pitch within ±4 frequency taps.

3.3. Results

The accuracy of the Pitch Activity Detector on the Test Set was 99.75%. The accuracy of the Pitch Estimator on the Test Set was 96.92%.

4. Discussion and Conclusions

In this paper, we presented DeepPitch, a Deep Convolutional Neural Network architecture for monophonic pitch estimation operating on log-frequency spectral inputs. DeepPitch is designed to handle a wide range of input signal types and effectively matches the capabilities of the Human Auditory Pathway. To extend DeepPitch to the general-purpose Polyphonic case, further development is necessary.

References

  1. Rachel M Bittner, Justin Salamon, Mike Tierney, Matthias Mauch, Chris Cannam, and Juan Pablo Bello, “Medleydb: A multitrack dataset for annotation-intensive mir research.”
  2. Juan Bosch and Emilia Gómez, “Melody extraction in symphonic classical music: a comparative study.”
  3. Matthias Mauch, Chris Cannam, et al., “Computer-aided melody note transcription using the tony software.”
  4. Maria Luisa Zubizarreta, “Prosody, focus, and word order.”
  5. Bregman, A., “Auditory Scene Analysis.”
  6. Goldstein, J. (1973). “An optimum processor theory for pitch of complex tones.”
  7. Terhardt, E. (1974). “Pitch consonance and harmony.”
  8. Michael Noll, “Cepstrum pitch determination.”
  9. John Dubnowski, Ronald Schafer, and Lawrence Rabiner, “Real-time digital hardware pitch detector.”
  10. Myron Ross et al., “Average magnitude difference function pitch extractor.”
  11. David Talkin, “A robust algorithm for pitch tracking (rapt).”
  12. Paul Boersma, “Accurate short-term analysis of the fundamental frequency.”
  13. Alain De Cheveigné and Hideki Kawahara, “Yin: a fundamental frequency estimator.”
  14. Matthias Mauch and Simon Dixon, “pYIN: A fundamental frequency estimator using probabilistic threshold distributions.”
  15. L. Watts, "Commercializing Auditory Neuroscience."
  16. Jong Wook Kim et al., “CREPE: A Convolutional Representation for Pitch Estimation.”
  17. Nitish Srivastava et al., “Dropout: a simple way to prevent overfitting.”
  18. Diederik Kingma and Jimmy Ba, “Adam: A method for stochastic optimization.”
  19. TensorFlow: https://www.tensorflow.org/.
  20. WaveSurfer: https://sourceforge.net/projects/wavesurfer/.
  21. L. Watts, "Reverse-Engineering the Human Auditory Pathway."