Field Notes from Delivery · Issue 03
Sounds like a Five-minute change: Red, Green, Refactor
ISSUE 03: This Week: TECHNICAL NOTES · TEST-DRIVEN DEVELOPMENT Your team just got asked to add a free-shipping rule. Follow along as you build it one failing test at a time, the way Kent Beck and Robert Martin intended.
By Vikas Agarwal ·

ISSUE 03: This Week: TECHNICAL NOTES · TEST-DRIVEN DEVELOPMENT
Your team just got asked to add a free-shipping rule. Follow along as you build it one failing test at a time, the way Kent Beck and Robert Martin intended.
Here's the scenario: your team just got a one-line request from the business. Give customers free shipping on any order over $50. Sounds like a five-minute change. Follow along and build it with TDD, and you'll see why it isn't.
Robert C. Martin wrote the discipline down as three rules, stricter than most people who say they practice TDD follow:
- You may not write production code until you have written a failing unit test.
- You may not write more of a unit test than is sufficient to fail, and not compiling counts as failing.
- You may not write more production code than is sufficient to pass the one currently failing test.
The third rule means writing the minimum every time, even when the minimum looks stupid.
Cycle 1: start with the case you're most sure of
RED
You pick the easiest case to state first: a $30 order clearly shouldn't get free shipping. Write that down as a test before writing the code.
1def test_order_under_50_does_not_get_free_shipping():2 assert qualifies_for_free_shipping(order_total=30) is False
Run it, and it fails right away. qualifies_for_free_shipping doesn't exist yet, and an undefined function counts as red.
GREEN
Now write the smallest thing that makes it pass. The one-line code:
1def qualifies_for_free_shipping(order_total):2 return False
This function will always return no. You will do it on purpose. And it satisfies the one test that you wrote earlier.
Cycle 2: one more test and the shortcut breaks
RED
Add the most important test: an order of exactly $50 should qualify.
1def test_order_of_50_or_more_gets_free_shipping():2 assert qualifies_for_free_shipping(order_total=50) is True
GREEN
return False can no longer pass both tests, so you will now have to right the real rule.
1def qualifies_for_free_shipping(order_total):2 return order_total >= 50
According to Kent Beck, this is called triangulation: a second test, chosen specifically to make the shortcut impossible, so the code has to generalize instead of piling up special cases.
Cycle 3: the requirement you didn't know about yet
RED
A week later, someone from operations added the requirement that heavy orders don't qualify, even over $50, because of profitability: shipping costs eat the margin.
1def test_heavy_order_does_not_qualify_even_over_50():2 assert qualifies_for_free_shipping(order_total=80, weight_kg=25) is False
GREEN
1def qualifies_for_free_shipping(order_total, weight_kg=0):2 return order_total >= 50 and weight_kg < 20
REFACTOR
In the code above, we wrote two conditions on one line. What if there are three? It will start getting complex. So you pull each rule into its own named function.
1def qualifies_for_free_shipping(order_total, weight_kg=0):2 return meetsminimum(order_total) and nottoo_heavy(weight_kg)
def meetsminimum(order_total):
return order_total >= 50
def nottoo_heavy(weight_kg):
return weight_kg < 20
Every test from cycles 1 and 2 still passes.
Every new rule from here will follow the same standard. One red test per rule. The smallest code that turns it green. A refactor only once the function gets hard to read.
Why should we take this trouble?
A 2008 study out of Microsoft Research tracked four industrial teams, three at Microsoft and one at IBM, who adopted this discipline on production codebases and measured defect density against each team's own prior baseline. The figures show fewer defects across every team:
IBM device drivers: 40%
Microsoft Windows: 62%
Microsoft MSN: 76%
Microsoft Visual Studio: 91%
Fewer defects across every team. The cost showed up on the other side of the ledger: the same researchers logged 15 to 35 percent more time spent on initial development. A team weighing a 40 to 91 percent drop in defects against roughly a quarter more development time usually doesn't need much convincing once they've seen what the alternative costs in production.
What is not TDD?
TDD run this way has nothing to do with a coverage target, and chasing one defeats the point. The tests exist to force a design decision when someone is best qualified to make it, before the code exists to defend it. A test written after the code, describing what the code already does, serves a different purpose: documentation more than design pressure. Rule one matters for a reason: the test has to fail first, or it was never driving anything.
I take engineering teams through this level of strictness when they've read about TDD. IT means minimum code, every cycle, no exceptions. Most break their own rule inside the first hour. Watching exactly where and why is usually more useful than the training itself.
Past the Training Workshop
Rule three breaks down the moment the management decides that it is a five-minute change. With tightened deadlines, nobody holds the discipline and start customizing TDD. Coaching addresses that gap. This is the culture change that we aim for.
Our custom corporate trainings are built to fill this gap. We assess why a team's stated practices and its behavior under pressure diverge, then build change that survives past the workshop.
A different piece of the problem shows up when TDD alone can't reach it, the cross-team and dev-to-ops handoffs that make "just write the test first" feel impossible when three other teams are blocking the work anyway. Delivery Without Silos exists to untangle exactly that.
Our offering, “Improving Ways of Working,” fits organizations past the training stage that are still watching the discipline slip in their teams. We build a diagnostic engagement to find the root cause of delivery slowdown. It could be a blocker in technical practice, team structure, or the incentives sitting on top of both.
For more details, visit www.provcraft.com
Source for the defect-density figures: Nagappan, Maximilien, Bhat & Williams, "Realizing Quality Improvement Through Test-Driven Development: Results and Experiences of Four Industrial Teams," Empirical Software Engineering, Vol. 13, 2008. The Three Laws of TDD are Robert C. Martin's; the fake-it/triangulation technique is Kent Beck's, from Test-Driven Development: By Example.
#TDD #TestDrivenDevelopment #SoftwareEngineering #CleanCode #EngineeringExcellence #AgileCoaching #SoftwareCraftsmanship #DevOps #CodeQuality #CorporateTraining #TechLeadership #RedGreenRefactor #ContinuousImprovement #ProvCraft
This issue was also published on LinkedIn. Join the conversation there or subscribe to Field Notes from Delivery.