Speech • Cognition • Technology

SPEAK Lab @ Waseda

Speech, Perception, Engineering, & AI Knowledge

We study how speech conveys meaning beyond words, and how insights from speech science can support more transparent, human-centered speech technologies.

Mission

Connecting speech phenomena, cognition, and speech technology.

SPEAK Lab brings linguistic analysis, cognitive explanation, empirical methods, and AI-informed engineering together to better understand human speech and improve speech agents.

To advance theory and practice at the intersection of speech phenomena, cognition, and speech technology.

Human speech carries rich information through prosody, timing, fluency, voice quality, and interactional feedback. We investigate these phenomena from linguistic and cognitive perspectives, then apply the findings to speech technologies that can process and produce speech more appropriately.

◌ Prosody◌ Fluency◌ Perception◌ Speech AI◌ Explainability

Speech as signal

We examine acoustic, temporal, prosodic, and fluency-oriented features that shape interpretation.

Speech as cognition

We ask how speech behavior reflects planning, perception, learning, and interaction.

Speech as system

We use empirical evidence to inform perceptive and productive speech-technology pipelines.

Speech as interaction

We develop and evaluate technologies that respond to human speech in transparent, useful ways.

Projects

Current research directions

Ongoing projects examine listeners, synthetic speech, second-language speech, and real-time feedback in human-computer communication.

LEDJ

Listener effects of disfluencies in Japanese

Testing how disfluencies shape listeners' interpretation of ongoing speech in Japanese.

PHASE

Prosodic Hesitation Analysis in Synthetic Environments

Exploring how AI agents handle hesitation and how closely synthetic behavior resembles human behavior.

L2 Pitch Accent

Second-language pitch accent in Japanese

Investigating where pitch-accent difficulties arise for nonnative speakers and how ASR systems might address them.

Fluidity

Real-time feedback in second-language speech

Developing feedback systems in which on-screen avatars respond to features of a speaker's fluency.

People

Researchers and collaborators

SPEAK Lab is based at Waseda University's Nishi-Waseda Campus and works across speech science, perceptual computing, neurolinguistics, language acquisition, and speech technology.

Ralph L. Rose, Ph.D.

Head researcher • Center for English Language Education in Science and Engineering (CELESE), Waseda University Faculty of Science and Engineering

Research interests include speech fluency, disfluency, prosody, corpus-based speech analysis, and the role of AI in language and speech technologies.

Selected works

Recent and representative publications

The first four items are shown by default. Use the button to show the complete publication list from the overview document.

From corpus studies of hesitation to speech-language models and AI-generated speech.

The lab's work combines linguistic description, perception experiments, speech technology evaluation, and ethical reflection on AI-generated speech.

Accepted

Orthography and pitch accent in Japanese TTS

Testing whether Japanese text-to-speech systems are sensitive to semantic information when realizing pitch accent.

Yang, Y. and Rose, R.L. (Accepted for presentation) Orthography Over Meaning? Testing Semantic Sensitivity in Japanese TTS Pitch Accent Realisation.
Accepted

Epistemic stance and LLM annotation

Comparing linguistic features, automatic annotation, and human judgment in the analysis of epistemic stance.

Rose, R.L., Sugawara, A. and Yang, Y. (Accepted for presentation) Do LLMs Understand Epistemic Stance? Evidence from Linguistic Features, Automatic Annotation, and Human Judgment.
Accepted

Disfluencies and clause-boundary detection

Showing how transcript representation matters for LLM-driven analysis of spoken-language corpora.

Rose, R.L. (Accepted for presentation) Disfluencies Matter: Transcript Representation and Large Language Model-driven Clause-Boundary Detection in a Speech Corpus.
Accepted

Clause boundaries and disfluency perception

Examining clause-boundary effects in disfluency perception using evidence from a speech language model.

Rose, R.L., Sugawara, A., and Yang, Y. (Accepted for publication) Clause boundary effects in disfluency perception: Evidence from a speech language model.
Contact

Interested in joining or collaborating?

Contact SPEAK Lab about collaboration, graduate research, speech corpus work, and human-centered speech technology projects.

Get in touch

Research conversations are welcome.

SPEAK Lab welcomes conversations with students and researchers interested in speech science, cognition, speech technology, and explainable human-centered AI systems.

rose@waseda.jp