arXiv cs.CL→ original

SAMPA: Whisper-Based System for Portuguese Speech Prosody Boundary Detection

Researchers developed SAMPA — an adaptation of Whisper large-v3 for Portuguese speech. The system automatically identifies prosodic boundaries between speech units, trained on recordings from the NURC-SP dataset, and demonstrated strong results: F1=0.731 on the main test and F1=0.796 on the MuPe-Diversidades dataset.

AI-processed from arXiv cs.CL; edited by Hamidun News
SAMPA: Whisper-Based System for Portuguese Speech Prosody Boundary Detection
Source: arXiv cs.CL. Collage: Hamidun News.
◐ Listen to article

Researchers developed SAMPA — a system based on the Whisper model for automatic detection of prosodic boundaries (points of separation between speech units) in Brazilian Portuguese speech. On the standard test, the system showed F1=0.731 accuracy, and in cross-dataset testing on MuPe-Diversidades achieved F1=0.796.

How SAMPA Works

The system is built on Whisper large-v3 — a multilingual speech-to-text model from OpenAI. The authors fine-tuned it on Portuguese recordings from the NURC-SP dataset (an archive of Brazilian Portuguese oral speech) and trained the model to insert special markers at prosodic boundary locations. The system was trained on recordings with different speakers, different speech styles, and different recording conditions, which helped it handle new examples as well.

Key aspects of the approach:

  • Base model: Whisper large-v3 from OpenAI
  • Fine-tuning: NURC-SP dataset with manually annotated recordings
  • Accuracy on main test: F1=0.731
  • Accuracy on independent MuPe-Diversidades dataset: F1=0.796
  • The model uses morphosyntactic, semantic, and prosodic signals

Why This Matters for Portuguese

For English, prosody processing is well-developed due to abundant annotated data and years of research. Portuguese, however, has historically relied mainly on rules and traditional machine learning methods — this is less accurate and flexible than neural network approaches.

SAMPA demonstrates that a large pre-trained model can be effectively adapted to another language using a relatively small set of annotated recordings. This opens the door to similar solutions for other languages beyond English — Russian, Spanish, Mandarin, and others can obtain comparable systems through Whisper adaptation.

What This Means

Correct detection of prosodic boundaries is important for practical Portuguese speech processing: natural speech synthesis sounds more convincing when pauses and intonation are correctly placed; automatic text annotation becomes possible without manual editing; analysis of oral documents and speech archives is accelerated.

The research demonstrates that the transfer learning principle works effectively — instead of training a model from scratch, the authors adapted an existing Whisper, which is significantly faster and requires less data. This could become a template for developing speech processing in other languages and regions.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Need AI working inside your business — not just in your newsfeed?

I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).

What do you think?
Loading comments…