A 2025 systematic review found an alarming range in AI-based speech recognition performance: while some controlled dictation settings achieved word-error rates under 10%, this figure climbed to over 50% in conversational and multi-speaker audio, showing significant problems handling specialized terminology and accented speech. [1]
Automated transcription may be good enough when the text is used only as a searchable record or a quick way to review an interview. But the standard changes when the transcription becomes part of the research data itself. Errors can then affect coding, analysis, quotations, and ultimately, published findings.
What makes AI errors particularly difficult to identify is that an AI transcription can read perfectly well yet still be wrong. It might assign a statement to the wrong speaker, miss a key word, change a technical term, or alter a quotation. Errors like these alter findings and could dangerously skew the researcher’s analysis. And AI transcription studies don't account for them in their accuracy scores.
The key question is not whether AI can create readable text. It is whether the transcription is accurate enough for how the research team plans to use it. Can they rely on it?
1. A 2024 research methods paper found that automated transcription can be useful for qualitative research in some cases. But its choices about speaker diarization (who said what), grammar, background noise, and speech quality are unreliable. Also, ASR has an especially difficult time separating speaker exchanges (starts/stops). [2]
2. A 2025 health research study compared professional human transcription with two AI transcription platforms. AI was faster and less costly, but it had problems with speaker identification, punctuation, cultural terms, and accented speech. The researchers stressed the need for human review. [3]
3. A 2024 study tested four ASR systems for racial bias in patient-nurse conversations. All four ASR systems were less accurate for Black patients than for White patients. Even the best system had a median word-error rate of 39%. [4]
The word-error rate tells how many words were wrong. It does not tell us which words were wrong or whether an error changed the meaning.
That matters in qualitative research. Missing the word "not" can reverse an answer. Assigning a statement to the wrong person can change the findings. A wrong medication, diagnosis, or technical term can affect coding and analysis. Confusing hypertensive and hypotensive can be critical.
These risks matter most when researchers use transcriptions for coding, compare responses, publish direct quotations, or work with sensitive or complex recordings.
For these uses, the better question is not, "Did the software understand the recording?" It is, "Can we trust this text as research data?"
Automated transcription systems struggle to determine when one person stops speaking and another begins, so they often misattribute participant feedback. This process is called speaker diarization, or more commonly referred to as speaker identification. Errors are especially prevalent when people interrupt each other, speak briefly, or have similar-sounding voices. [2][3]
A sentence can contain every correct word and still be bad research data if it is assigned to the wrong person.
For focus groups and other recordings with several speakers, researchers should check speaker labels as a separate quality step.
The 2025 review found ongoing errors with specialized terms and accented speech. [1] This is important when interviews include medical, scientific, or other technical language.
For more on this issue, see Research Transcriptions' article on industry-specific language and AI transcription: Why AI-Powered Speech Recognition Fails at Industry-Specific Language
This is especially important for VA researchers. VA guidance updated on July 22, 2026, says VA sensitive data may be used only with tools that have VA Authority to Operate for that purpose. The guidance includes PHI and PII as sensitive data. It also says VA staff remain responsible for work or decisions produced with generative AI. [6]
This does not mean VA researchers cannot use AI. VA is expanding access to approved AI tools that meet its security and privacy rules. The specific tool, type of data, approval status, and planned use all matter. [6]
Learn about Research Transcriptions' SOC-2-certified security and compliance.
Automated transcription can be a solid choice when minor errors are unlikely to affect the work and the recording is clear.
Examples include:
In these cases, automated transcription can save time and money.
Human transcription adds a level of review that automated transcription cannot provide. A trained transcriptionist can use context, identify speakers, resolve unclear speech, and check the final transcription for accuracy.
The choice does not have to be AI or human transcription in every case. Researchers can choose the method based on how they will use the transcription.
Start by deciding how important accuracy and security are, and how much time you have to review and correct the transcription before choosing a tool.
Plenty of research shows both the benefits and the limitations of AI transcription. It can be fast and cheap, but accuracy can change based on the recording, the speakers, accents, and technical language. And in the end, “cheap” can end up being very expensive. [1][2][3][4]
As the importance, difficulty, or sensitivity of the recording increases, human review should increase too. The key question is not how fast and cheap the transcription can be created. It is whether the research team can trust the text to reflect what participants actually said.
[1] Ng JJW, Wang E, Zhou X, et al. Evaluating the performance of artificial intelligence-based speech recognition for clinical documentation: a systematic review. BMC Medical Informatics and Decision Making. 2025;25:236. doi:10.1186/s12911-025-03061-0. https://pmc.ncbi.nlm.nih.gov/articles/PMC12220090/
[2] Eftekhari H. Transcribing in the digital age: qualitative research practice utilizing intelligent speech recognition technology. European Journal of Cardiovascular Nursing. 2024;23(5):553-560. doi:10.1093/eurjcn/zvae013. https://pmc.ncbi.nlm.nih.gov/articles/PMC11334016/
[3] Kabir SMA, Ali F, Sulaiman-Hill R. A comparative assessment of AI and manual transcription quality in health data: insights from field observations. New Zealand Medical Journal. 2025;138(1625):35-43. doi:10.26635/6965.7024. https://pubmed.ncbi.nlm.nih.gov/41197094/
[4] Zolnoori M, Vergez S, Xu Z, et al. Decoding disparities: evaluating automatic speech recognition system performance in transcribing Black and White patient verbal communication with nurses in home healthcare. JAMIA Open. 2024;7(4):ooae130. doi:10.1093/jamiaopen/ooae130. https://pmc.ncbi.nlm.nih.gov/articles/PMC11631515/
[5] Samuel G, Wassenaar D. Joint Editorial: Informed Consent and AI transcription of Qualitative Data. Journal of Empirical Research on Human Research Ethics. 2025;20(1-2):3-5. doi:10.1177/15562646241296712. https://pmc.ncbi.nlm.nih.gov/articles/PMC12048736/
[6] U.S. Department of Veterans Affairs. Guidance for Generative AI Use at VA. Updated July 22, 2026. https://department.va.gov/ai/guidance-for-generative-ai-use-at-va/
Research transcriptions. Why AI-powered Speech Recognition Fails at Industry-Specific Language. https://blog.researchtranscriptions.com/ai-transcription-industry-language_a07
Read About Research Transcriptions' Confidentiality
Read about Research Transcriptions’ Academic Research Transcription Service
Research transcriptions. Medical Research Transcription Service