Common methods of measuring how people understand programming code are being analyzed in a new way by researchers at New Jersey Institute of Technology, bringing a technique from social science into the world of computer science.
Their goals are to correct what they see as flawed approaches in the study of code comprehension and ultimately to teach programmers to write code that is more easily understood from the start.
“We did an extensive literature review, and we were surprised because we found out apparently the whole field of program comprehension, after more than 40 years of publishing papers and studying on this matter, apparently it's been built on shaky foundations … There’s no explicit standard definitions and no reliable ways of measuring it,” said Erfan Arvan, a fourth-year doctoral student in computer science at NJIT.
It turns out that simply asking someone to explain what output the code would produce is the best indicator of whether they understand it, according to a study of how experienced developers work in the real world. Measuring how long someone took to determine that output is the next-best way, Arvan said.
“LLMs are generating a lot of code but we're not sure about a lot of aspects,” he said, such as whether it’s reliable, secure, understandable and verifiable. “Our idea was to see whether we can propose and create a tool or a system that can automatically detect whether this piece of code is understandable to humans or not. This is important because we already know that software engineers and programmers spend a considerable amount of their time trying to understand code.”
Arvan, along with NJIT assistant professor Martin Kellogg and collaborators at William & Mary, published their work On the Reliability of Code Comprehension Proxies and won a Distinguished Paper Award at the IEEE/ACM Automated Software Engineering conference held in Munich this fall.
Emphasizing a code snippet’s bottom line — what it does — is far more important than focusing on syntax, Arvan said. It’s similar to how linguists would observe that in everyday life, non-native speakers are judged by how well they get their message across, not by their grammar or spelling.
To start, Arvan and peers used the Delphi method which they noted is “widely used for expert consensus under conditions of uncertainty in other domains such as medicine and national-security forecasting, but to our knowledge, this is its first application to code comprehension research.”
“Over multiple rounds, participants ranked the [Java] snippets by comprehension difficulty and iteratively resolved disagreements through written feedback and discussion,” they stated, to determine a basis for which code was good and which was not. Next, they presented dozens of students with several common evaluation methods. These included syntax, timed input/output questions and self-evaluations about code rating.
Another motivation for their work was that many software development environments apply a code quality measurement based on what Kellogg said is a discredited concept called cyclomatic complexity, involving a score based on how many decisions a program can make. This is indicative of a wider problem, which is that too often the code is evaluated based merely on what automation tools and development environments can prove, instead of what matters to real-world engineers. Their paper didn’t examine that particular issue, but Arvan and Kellogg hope their new work can motivate those who build software development tools to modernize such approaches.
Looking forward, “A future direction of our plans is to revisit all of the prior studies that use one of these measurements we test, and we want to show that to the community. We want to review all of the papers in light of our results, so that we can see which results can be trusted, which results cannot be trusted, which papers should be redone and reconduct their studies so the results will be reliable,” Arvan said. Collaborators at William & Mary will study the possibility of replacing human participants in these studies with LLM agents, he added.