|国家预印本平台
首页|Do Code LLMs Do Static Analysis?

Do Code LLMs Do Static Analysis?

Do Code LLMs Do Static Analysis?

来源:Arxiv_logoArxiv
英文摘要

This paper investigates code LLMs' capability of static analysis during code intelligence tasks such as code summarization and generation. Code LLMs are now household names for their abilities to do some programming tasks that have heretofore required people. The process that people follow to do programming tasks has long been understood to require static analysis. For example, human programmers navigate the call graph of large programs to comprehend the different parts of those programs. Education in programming includes static analysis under the assumption that better static analysis skills beget better programming. Yet while popular culture is replete with anthropomorphic references such as LLM "reasoning", in fact code LLMs could exhibit a wholly alien thought process to humans. This paper studies the specific question of static analysis by code LLMs. We use three different static analysis tasks (callgraph generation, AST generation, and dataflow generation) and three different code intelligence tasks (code generation, summarization, and translation) with two different open-source models (Gemini and GPT-4o) and closed-source models (CodeLlaMA and Jam) as our experiments. We found that LLMs show poor performance on static analysis tasks and that pretraining on the static analysis tasks does not generalize to better performance on the code intelligence tasks.

Chia-Yi Su、Collin McMillan

计算技术、计算机技术

Chia-Yi Su,Collin McMillan.Do Code LLMs Do Static Analysis?[EB/OL].(2025-05-17)[2025-06-17].https://arxiv.org/abs/2505.12118.点此复制

评论