Archive for Computational linguistics

Phil Resnik's ACL keynote: A new balancing act

On July 5, Phil Resnik delivered a keynote address at the 64th Annual Meeting of the Association for Computational Linguistics, with the title "A New Balancing Act: Reflections on the Relationship between Computational Linguistics and AI":

In this talk I argued that the field of computational linguistics – a term that includes NLP as its engineering-research subdiscipline – is experiencing a “success catastrophe”. The commercial success of LLM-based AI has thrown three key aspects of our research community out of balance. Here are three balancing acts we face:

First: Like any research community, we can recognize and take advantage of the knowledge obtained in earlier generations of work; we can also lean into new approaches.

Second: We can focus on language as language, which is to say, the properties of language that make it distinctive and human; we can also treat language as an input/output modality for AI systems.

Third: We can emphasize our role as a research community, where our primary purpose is to contribute to the stock of human knowledge; we can also emphasize our role in making sure that our young members have a path forward to get jobs – particularly jobs in industry since the path into academia is never a sure bet and for many of them industry is the goal.

In each of these pairings, the central importance of the former has given way to the overwhelming dominance of the latter.

Read the rest of this entry »

Comments (6)

Essence of meaning

Comments (6)

Word frequencies in LOTR vs. Dickens

Following up on "Meadow writing", I thought it might be interesting to look at LOTR-associated word frequencies, using the the "weighted log-odds-ratio, informative dirichlet prior" algorithm Monroe, Colaresi, and Quinn 2009, "Fightin' Words", as discussed in seven previous LLOG posts. In particular, I thought I'd compare The Fellowship of the Ring to 16 of Charles Dickens' works.

Read the rest of this entry »

Comments (3)

What's (still) wrong with text-to-speech?

Text-To-Speech technology has improved enormously over the decades — but there's still some headroom, as a friend has recently underlined for me. He observes that when The Economist magazine first publishes a piece online, it appears with a AI-read audio, and then later with a human-read version:

The rhythm/prosody/pitch (I'm not exactly sure which – all three?) is the same in nearly every sentence and even clause. This high-then-falling pattern is fine in one sentence, but repeated 50 times in a row is awful.

Later, those pieces that make it into the print edition get their own, human-read version. So voilà, you have a perfect before-and-after.

Read the rest of this entry »

Comments (8)

Unifying Arabic topolects through AI

Meet Habibi – the Chinese AI uniting 20 Arabic dialects in a Middle East first
Lead author says there are many differences between Arabic dialects and Modern Standard Arabic, which is used in official circumstances
Zhao Ziwen, SCMP, 28 Feb 2026

The paper that presents this new model is called “Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis”. It was published last month on arXiv, an open-access repository that is not peer-reviewed.  I will be interested to hear what Language Log readers think of its prospects.

Read the rest of this entry »

Comments (4)

Latent trees

There's been some buzz recently about how syntactic structures are implicit in Large Language Models — most recently, the Liu et al. paper noted yesterday by Victor, and an accepted ms by Futrell and Mahowald at Behavioral and Brain Sciences, "How Linguistics Learned to Stop Worrying and Love the Language Models". Futrell and Mahowald recognize something that Liu et al. mostly ignore, namely that constituent structure is obviously implicit in statistical patterns of sequential data, at least if the sequences were generated by a constituency-sensitive process — and that algorithms taking advantage of that fact have been Out There for 70 years or more.

Read the rest of this entry »

Comments (2)

LLMs and tree-structuring

"Active Use of Latent Tree-Structured Sentence Representation in Humans and Large Language Models." Liu, Wei et al. Nature Human Behaviour (September 10, 2025).

Abstract

Understanding how sentences are represented in the human brain, as well as in large language models (LLMs), poses a substantial challenge for cognitive science. Here we develop a one-shot learning task to investigate whether humans and LLMs encode tree-structured constituents within sentences. Participants (total N = 372, native Chinese or English speakers, and bilingual in Chinese and English) and LLMs (for example, ChatGPT) were asked to infer which words should be deleted from a sentence. Both groups tend to delete constituents, instead of non-constituent word strings, following rules specific to Chinese and English, respectively. The results cannot be explained by models that rely only on word properties and word positions. Crucially, based on word strings deleted by either humans or LLMs, the underlying constituency tree structure can be successfully reconstructed. Altogether, these results demonstrate that latent tree-structured sentence representations emerge in both humans and LLMs.

Read the rest of this entry »

Comments (7)

Linguistics Olympiad

Taiwan hosts its 1st International Linguistics Olympiad
Nearly 400 people competing at National Taiwan University
Kelvin Chen, Taiwan News (7/21/25)

I wonder whether any Language Log readers have heard of the International Linguistics Olympiad or may even have participated in one of the Olympiads that have been held in at least 23 different countries since its founding in 2003.  Because it has an interesting history and purpose, before telling about what is happening in Taiwan right now (July 21-26, 2025), I'll give a brief sketch of the origins and aims of the Olympiad:

Read the rest of this entry »

Comments (7)

Computational phylogeny of Indo-European

Alexei S. Kassian and George Starostin, "Do 'language trees with sampled ancestors' really support a 'hybrid model' for the origin of Indo-European? Thoughts on the most recent attempt at yet another IE phylogeny".  Humanities and Social Sciences Communications, 12, no. 682 (May 16, 2025).

Abstract

In this paper, we present a brief critical analysis of the data, methodology, and results of the most recent publication on the computational phylogeny of the Indo-European family (Heggarty et al. 2023), comparing them to previous efforts in this area carried out by (roughly) the same team of scholars (informally designated as the “New Zealand school”), as well as concurrent research by scholars belonging to the “Moscow school” of historical linguistics. We show that the general quality of the lexical data used as the basis for classification has significantly improved from earlier studies, reflecting a more careful curation process on the part of qualified historical linguists involved in the project; however, there remain serious issues when it comes to marking cognation between different characters, such as failure (in many cases) to distinguish between true cognacy and areal diffusion and the inability to take into account the influence of the so-called derivational drift (independent morphological formations from the same root in languages belonging to different branches). Considering that both the topological features of the resulting consensus tree and the established datings contradict historical evidence in several major aspects, these shortcomings may partially be responsible for the results. Our principal conclusion is that the correlation between the number of included languages and the size of the list may simply be insufficient for a guaranteed robust topology; either the list should be drastically expanded (not a realistic option for various practical reasons) or the number of compared taxa be reduced, possibly by means of using intermediate reconstructions for ancestral stages instead of multiple languages (the principle advocated by the Moscow school).

Read the rest of this entry »

Comments (5)

Unicode CJK Unified Ideographs Extension J and the nature of the sinographic writing system

Submitted by Charles Belov:

I've been browsing through the proposed Unicode 17 changes, currently undergoing a comment period, with interest. While I don't have the knowledge to intelligently comment on the proposals, it's good to see that they are actively improving language access.

I'm puzzled that some new characters have been added to the existing Unicode CJK Unified Ideographs Extension C (6 characters) and Unicode CJK Unified Ideographs Extension E (12 characters) rather than added to a new extension. But the most interesting is the apparently brand-new Unicode CJK Unified Ideographs Extension J, with over 4,000 added characters.

Read the rest of this entry »

Comments (32)

A new voice morphing application

Over the years, we've documented various applications of voice morphing technology besides the malicious creation of "deep fake" audio clips. Here's a new one: Amrit Dillon, "AI erases call centre staff’s Indian accents", The Times 3/2/2025:

A French company which operates the largest number of call centres in the world is using artificial intelligence to soften Indian accents in real time to make customer conversations easier and shorter.

Teleperformance said that it was sometimes difficult for customers calling call centres in India — and the Philippines — to understand workers’ accents, leading to frustration and longer than necessary calls.

“When you have an Indian agent on the line, sometimes it’s hard to hear, to understand,” Thomas Mackenbrock, the company’s deputy chief executive, told Bloomberg News. “The technology can neutralise the accent of the Indian speaker with zero latency. This creates more intimacy, increases customer satisfaction, and reduces the average handling time. It is a win-win for both parties.”

The software, called “accent translation”, has been developed by Sanas, a start-up based in Palo Alto, California.

Read the rest of this entry »

Comments (19)

Remaining problems with TTS

(…and with the New York Department of Environmental Conservation…)

Like many other online text sites, the New York Times now offers synthetic text-to-speech readings for (most of) its stories. TTS quality has improved enormously since the 1980s, when I worked with Bill Dunn from Dow Jones Information Services on (the idea of) a pre-internet version of digital news delivery, including synthesized audio versions. (See "Thanks, Bill Dunn!", 8/6/2009, for a bit more of the story.)

And this morning, while doing some brainless form checking, I listened to the audio version of Victor Mather and Jesus Jiménez, "After 7 Years, P’Nut the Squirrel Is Taken Away and Then Put Down", NYT 11/1/2024, which starts this way:

P’Nut, a pet squirrel with a popular Instagram page, was seized by state government officials on Wednesday in Pine City, N.Y., and later euthanized to test for rabies.

Read the rest of this entry »

Comments (10)

Psychotic Whisper

Whisper is a widely-used speech-to-text system from OpenAI — and it turns out that generative AI's hallucination problem afflicts Whisper to a surprisingly serious extent, as documented by Allison Koenecke, Anna Seo Gyeong Choi, Katelyn X. Mei, Hilke Schellmann, and Mona Sloane,"Careless Whisper: Speech-to-Text Hallucination Harms", In The 2024 ACM Conference on Fairness, Accountability, and Transparency,  2024:

Abstract: Speech-to-text services aim to transcribe input audio as accurately as possible. They increasingly play a role in everyday life, for example in personal voice assistants or in customer-company interactions. We evaluate Open AI’s Whisper, a state-of-the-art automated speech recognition service outperforming industry competitors, as of 2023. While many of Whisper’s transcriptions were highly accurate, we find that roughly 1% of audio transcriptions contained entire hallucinated phrases or sentences which did not exist in any form in the underlying audio. We thematically analyze the Whisper-hallucinated content, finding that 38% of hallucinations include explicit harms such as perpetuating violence, making up inaccurate associations, or implying false authority. We then study why hallucinations occur by observing the disparities in hallucination rates between speakers with aphasia (who have a lowered ability to express themselves using speech and voice) and a control group. We find that hallucinations disproportionately occur for individuals who speak with longer shares of non-vocal durations—a common symptom of aphasia. We call on industry practitioners to ameliorate these language-model-based hallucinations in Whisper, and to raise awareness of potential biases amplified by hallucinations in downstream applications of speech-to-text models.

Read the rest of this entry »

Comments (12)