江汉学术 ›› 2025, Vol. 44 ›› Issue (4): 73-83.doi: 10.16388/j.cnki.cn42-1843/c.2025.04.007

• 语言学专栏_领域语言研究 • 上一篇    下一篇

大语言模型的语言能力评测研究:特征、路径和趋势

易保树1,倪传斌2   

  1. 1南京邮电大学 外国语学院,南京 210023;2南京师范大学 外国语学院,南京 210023
  • 收稿日期:2025-04-27 出版日期:2025-08-15 发布日期:2025-08-15
  • 作者简介:易保树,男,安徽巢湖人,南京邮电大学外国语学院副教授,博士,E-mail:yibaoshu111@163.com;倪传斌,男,湖北荆州人,南京师范大学外国语学院教授,博士生导师,E-mail:nichuanbin1985@126.net。
  • 基金资助:
    江苏省社会科学基金项目“句法孤岛加工和习得中工作记忆作用机制研究”(23YYB004);江苏高校外语教育“高质量发展背景下外语教学改革”专项课题“大语言模型在大学英语读写教学中的应用”(2024WYJG094)

Linguistic Competence Evaluation of Large Language Models:Feature,Approach and Trend

YI Baoshu1,NI Chuanbin2   

  1. 1School of Foreign Studies,Nanjing University of Posts and Telecommunications,Nanjing 210023;2School of Foreign Languages and Cultures,Nanjing Normal University,Nanjing 210023
  • Received:2025-04-27 Online:2025-08-15 Published:2025-08-15

摘要: 回顾大语言模型(Large Language Models,LLMs)语言能力发展研究,对比LLMs与人类语言学习特征的差异,从学习环境和机制、语言特异性泛化能力测量及语法能力测评多维度探讨LLMs的语言能力评测及其理论启示,可以发现:学习环境层面,LLMs凭借海量单模态文本输入实现高效统计泛化,而人类在生态效度更高的多模态交互中发展语言能力,二者形成互补性差异;针对语言天赋论的核心假设,通过消融实验、无监督和监督测试三类范式,揭示LLMs虽缺乏人类先验语法特异性,却能通过统计模式复现部分语法规则;语法能力测评表明,LLMs虽可习得表层句法结构,但对深层递归性、语义—句法接口等人类特异性特征的建模仍存在显著局限。同时,LLMs的涌现能力对刺激贫乏论与语言天赋假设构成双重挑战,并推动计算语言学与理论语言学、认知科学等领域的范式融合。未来大模型语言能力测评需聚焦语言形式与功能的认知解耦机制,探索跨学科方法论协同路径,以厘清LLMs语言能力边界。

关键词: 人工智能, 大语言模型, 语言能力, 语法能力, 语言习得, 句法加工

Abstract: After reviewing researches on the development of linguistic competence of Large Language Models(LLMs)and comparing the different characteristics between LLMs and human speech learning,this study explores the evaluation of LLMs’linguistic competence and its theoretical implications from multiple dimensions,including the learning environment and mechanism,the measurement of languagespecific generalization ability,and the assessment of grammatical competence. It can be found that:In terms of learning environment,LLMs achieve efficient statistical generalization with massive single-modal text input,while humans develop language capacity in multi-modal interactions with higher ecological validity; their differences are complementary. Regarding the core assumption of genetic theory of language,the results of ablation experiment,unsupervised and supervised tests reveal that although LLMs lack the prior grammatical specificity of humans,they can reproduce some grammatical rules through statistical models. The assessment of grammatical competence indicates that although LLMs can acquire surface syntactic structures,there are still significant limitations in modeling human-specific features such as deep recursion and semantic-syntactic interfaces. Meanwhile,the emergent ability of LLMs poses a dual challenge to the theory of stimulus scarcity and the genetic theory of language;it promotes the paradigm fusion of computational linguistics with theoretical linguistics,cognitive science and other fields. In the future, the assessment of LLMs’ language capabilities needs to focus on the cognitive decoupling mechanism between language forms and functions,so as to explore the collaborative approaches of interdisciplinary methodologies and clarify LLMs’language capability boundaries.

Key words: artificial intelligence (AI), Large Language Model (LLM), linguistic competence, grammatical competence, language acquisition, syntactic processing

中图分类号: