Ask a Mandarin speaker to say “bit” and “beat” in quick succession, and something curious happens. The vowels come out different, clearly different, but not necessarily in the way a native English speaker would make them different. English distinguishes these two vowels mostly through vowel quality: where the tongue sits in the mouth changes the sound’s spectral shape. Length matters too, but it’s the minor cue. Mandarin speakers often flip that hierarchy. For them, duration becomes the main signal and quality fades into the background. The vowels are still contrasted. They’re just contrasted using the wrong tool, at least by native standards.
This kind of mismatch has been documented for decades. What hasn’t been settled is a more basic question: when a learner’s ear and mouth disagree about which acoustic cue matters, are the two systems actually talking to each other at all? If someone leans on duration when listening to a vowel contrast, does that same lean show up when they try to reproduce it? Or do perception and production operate as separate systems that happen to produce roughly compatible outputs?

A new study1 out of National Yang Ming Chiao Tung University, led by Will Chih-Chao Chang and Yu-An Lu, gets at this question with more precision than most previous attempts, and the answer turns out to be genuinely interesting: it depends on which sound, and which acoustic dimension you’re asking about.









