Research article · Software engineering, economics & security

The work
after the code.

Software work continues after the first implementation. Testing, repair and the changes introduced by repair belong in the cost of delivering a dependable capability. Research on in-house development asks how that work should enter the productivity measure.

My 2011 paper with Tanveer Zia, A Quantitative Analysis into the Economics of Correcting Software Bugs, examines development and maintenance together. Its starting concern is the use of code volume as a measure of productive output.

Count the whole job

Consider a hypothetical comparison between two ways of implementing the same requirement. One produces a working demonstration quickly but needs repeated corrections. The other takes longer initially and requires less subsequent repair. A comparison that stops at the demonstration cannot establish which approach used fewer resources to deliver a dependable result.

The boundary of the measurement matters. Include the time spent understanding a failure, reproducing it, changing the implementation and checking the result. Record documentation and handover work too. Otherwise, an apparent gain in one stage can be a transfer of work to another person or a later budget.

Keep the historical evidence in view

The public manuscript describes 277 coding projects across 15 companies, with observations collected between June 2008 and December 2010. It records coding and debugging hours separately, together with code size and defects observed initially and through later patches. The projects involved in-house development in financial services, media and web development.

The study reports that repairs could introduce further defects, creating additional rounds of work. It also examines the rising effort required to find additional bugs. Those observations motivate an economic question about the full development cycle. Their reported rates and cost relationships belong to the historical sample; using them for another organisation requires fresh evidence.

Make a repair earn its completion

A useful repair record connects a reported failure to a specific change and a check that can detect the failure. It should also identify which neighbouring behaviours were examined. Closing an issue then has an evidential meaning: another person can see what was corrected and how the result was checked.

NIST’s Secure Software Development Framework, version 1.1, published in 2022, provides a complementary set of practices. PW.7 addresses source-code review and analysis. PW.8 addresses executable-code testing, including the recording and triage of findings. RV.3 adds investigation of root causes and changes intended to reduce recurrence.

These practices help organise the work. They do not supply a universal defect rate or a fixed number of testing hours. Their economic evaluation still needs a defined system, a stated testing scope and evidence about the consequences of failure.

Compare methods over a common horizon

A follow-on study could compare development methods against the same requirements and observe each through a declared maintenance period. It could record implementation, review and repair effort; distinguish defects by consequence; and track which changes created regressions. Differences in project difficulty and team experience would need to be accounted for.

The observation period should remain visible in the result. A defect discovered after the measurement closes can change the apparent cost of an earlier release. Equally, a method that finds more faults before delivery may have supplied better evidence rather than produced worse software.

This is a proposed comparison, not a report of new experimental findings. It carries the original question forward: what resources were required to deliver and sustain the intended behaviour? Answering that question gives testing and maintenance their proper place in an account of software productivity.