ABSTRACT:
Generative AI (GenAI) has advanced rapidly and made significant impacts. However, its effect on developers remains a topic of industry debate. Companies want to know whether GenAI can enhance developers’ coding performance, as an unclear understanding may put companies at a disadvantage. While the literature has begun addressing this issue, a formal understanding of GenAI’s impact remains incomplete. Moreover, existing findings are often short-term, fragmented, or lack explanatory mechanisms. To fill these gaps, we designed a multimethod research program comprising a longitudinal field study and a randomized controlled experiment. In Study 1, we collaborated with a global information technology organization and applied a difference-in-differences approach to over 27 weeks of proprietary data. In Study 2, we designed a randomized experiment involving 253 software developers. From these studies, we find that GenAI usage affects both developers’ coding quantity and quality. These effects, however, depend critically on how the tool is used. While reduced cognitive effort can be associated with diminished quality, interestingly, GenAI usage enables developers to produce higher-quality code with less cognitive effort. In this current study, we explain the paradoxical findings through cognitive load theory, showing that GenAI reduces extraneous load while preserving germane processing during ideation and debugging. Using a multimethod research design that integrates longitudinal field data with a randomized controlled experiment, we link observed performance effects to underlying cognitive mechanisms and usage strategies. We also offer guidance on effective usage styles and propose boundary conditions for realizing GenAI’s benefits in practice.
Key words and phrases: Generative AI, GenAI, GenAI coding, cognitive load, software developers, coding performance, difference-in-differences