首页|Do Code LLMs Do Static Analysis?

Do Code LLMs Do Static Analysis?

来源：

英文摘要

This paper investigates code LLMs' capability of static analysis during code intelligence tasks such as code summarization and generation. Code LLMs are now household names for their abilities to do some programming tasks that have heretofore required people. The process that people follow to do programming tasks has long been understood to require static analysis. For example, human programmers navigate the call graph of large programs to comprehend the different parts of those programs. Education in programming includes static analysis under the assumption that better static analysis skills beget better programming. Yet while popular culture is replete with anthropomorphic references such as LLM "reasoning", in fact code LLMs could exhibit a wholly alien thought process to humans. This paper studies the specific question of static analysis by code LLMs. We use three different static analysis tasks (callgraph generation, AST generation, and dataflow generation) and three different code intelligence tasks (code generation, summarization, and translation) with two different open-source models (Gemini and GPT-4o) and closed-source models (CodeLlaMA and Jam) as our experiments. We found that LLMs show poor performance on static analysis tasks and that pretraining on the static analysis tasks does not generalize to better performance on the code intelligence tasks.

作者：Chia-Yi Su、Collin McMillan

作者单位：

学科分类：计算技术、计算机技术

推荐引用：Chia-Yi Su,Collin McMillan.Do Code LLMs Do Static Analysis?[EB/OL].(2025-05-17)[2025-06-17].https://arxiv.org/abs/2505.12118.点此复制

Do Code LLMs Do Static Analysis?

Do Code LLMs Do Static Analysis?

评论