The Code Refactoring Advisor: Before-and-After Fixes With a Risk Score for Each One

Why this prompt matters
Developers now spend 23-42% of their work week dealing with technical debt and bad code — the Stripe Developer Coefficient study puts the global cost at $85 billion in lost productivity a year. The failure mode isn't not knowing code is bad; it's not knowing which fix is safe to make first. A one-standard-deviation rise in an organization's debt-to-code ratio corresponds to a 31% jump in defect density — refactoring the wrong part first, or all of it at once, is how technical debt cleanup itself introduces new bugs.
What we use it for
You've inherited a 400-line function that everyone on the team avoids touching because it 'works, mostly' and nobody remembers why it's structured the way it is. You need to clean it up before adding a new feature, but you don't have time to rewrite it from scratch and can't afford to break the parts that work.
Prompt
You are a senior software engineer with expertise in refactoring legacy code without introducing regressions. You are conservative by default — you flag risk honestly rather than being falsely reassuring. CONTEXT: Language/framework: [YOUR LANGUAGE, e.g. "TypeScript, Node.js, Express"] What this code does: [ONE-SENTENCE DESCRIPTION OF THE FUNCTION'S PURPOSE] Known problems (if any): [WHAT YOU ALREADY SUSPECT IS WRONG, e.g. "deeply nested conditionals, unclear variable names, does three unrelated things"] Test coverage: [DESCRIBE CURRENT TESTS, e.g. "one integration test covering the happy path only" or "none"] Constraints: [ANYTHING THAT CANNOT CHANGE, e.g. "the public function signature must stay identical, this is called from 40+ places"] TASK: Review the code below and identify 3-6 specific, independent refactoring opportunities. For EACH one: 1. Name the problem (e.g. "duplicated validation logic", "mixed abstraction levels", "unclear boolean flag parameter"). 2. Show the BEFORE code snippet (just the relevant lines, not the whole function). 3. Show the AFTER code snippet with the fix applied. 4. Assign a Risk Score (Low/Medium/High) based on: how much of the function's behavior the change touches, whether existing tests would catch a regression, and whether the change is purely structural (safe) or alters logic/behavior (riskier). 5. State what you'd want to verify manually or via a new test before merging this specific change. CONSTRAINTS: - Do NOT rewrite the entire function in one pass. Each refactor must be independently applicable — I need to be able to accept refactor #2 without necessarily taking #4. - Order your suggestions from LOWEST risk to HIGHEST risk, so I can bank easy wins first. - If the function is doing more than one job, say so explicitly and suggest a split — but still give me the low-risk internal cleanups first, since a full split is inherently higher-risk. - Do not suggest changes that alter the function's observable behavior unless you flag it as Medium or High risk with a clear explanation of what changes. - If test coverage is described as weak or absent, recommend which specific test(s) to write BEFORE attempting the Medium/High risk refactors, not after. OUTPUT FORMAT: ## Refactor 1: [Problem name] — Risk: [Low/Medium/High] **Before:** ``` [code] ``` **After:** ``` [code] ``` **Verify before merging:** [what to check or test] [repeat for each refactor, low to high risk] ## Summary [One paragraph: overall assessment of the function's health, and the recommended order of operations] CODE TO REVIEW: [PASTE YOUR FUNCTION OR CODE BLOCK HERE]
Result
Refactor 1: Duplicated validation logic — Risk: Low
Before:
if (!user.email || user.email.length === 0) { throw new Error('Email required'); }
if (!user.name || user.name.length === 0) { throw new Error('Name required'); }
if (!user.phone || user.phone.length === 0) { throw new Error('Phone required'); }After:
const requiredFields = ['email', 'name', 'phone'];
for (const field of requiredFields) {
if (!user[field]?.length) throw new Error(`${field} required`);
}Verify before merging: Confirm error message format isn't parsed elsewhere by exact string match — the field name capitalization changed slightly.
Refactor 2: Unclear boolean flag parameter — Risk: Low
Before:
function processOrder(order, flag) {
if (flag) { applyDiscount(order); }
...
}After:
function processOrder(order, { applyLoyaltyDiscount }) {
if (applyLoyaltyDiscount) { applyDiscount(order); }
...
}Verify before merging: Update all 6 call sites to pass a named object instead of a positional boolean — purely mechanical, no behavior change, but touches every caller.
Refactor 3: Deeply nested conditionals — Risk: Medium
Before:
if (user.active) {
if (user.subscription) {
if (user.subscription.tier === 'pro') {
// 15 more lines
}
}
}After:
if (!user.active) return;
if (!user.subscription) return;
if (user.subscription.tier !== 'pro') return;
// 15 more lines, now unindentedVerify before merging: Early returns change control flow if there's an else branch further down handling the negative case elsewhere in the original function — check for one before applying. Write a test for the non-pro-tier path specifically, since current coverage doesn't exercise it.
Refactor 4: Function does three unrelated things — Risk: High
Before: Single 400-line function handles validation, discount calculation, and email notification in sequence.
After: Split into validateOrder(), calculateDiscount(), and notifyCustomer(), called in sequence from a thin orchestrator function.
Verify before merging: This changes the function's shape enough that any code relying on partial execution (e.g. catching an error mid-function and still sending a notification) will behave differently. Do NOT attempt this split until Refactors 1-3 are merged and you have integration test coverage for all three code paths — this should be its own PR, reviewed separately.
Summary
This function is moderately unhealthy but not unsalvageable. Refactors 1-2 are safe, mechanical wins you can ship today. Refactor 3 needs one new test first but is still low-stakes. Refactor 4 (the actual split) is the real fix but should wait until the others are merged and stable — attempting all four at once is how a cleanup PR turns into a production incident.
Most refactoring prompts produce one giant rewrite of your function and hope it's correct. That's backwards for real production code — a single sweeping rewrite is exactly the kind of change that's hardest to review, hardest to test, and most likely to hide a regression inside a diff too large for anyone to properly check. This prompt does the opposite: it breaks the cleanup into independent, individually-mergeable pieces and ranks them by how much they could break.
Why risk scoring comes before code quality scoring
A refactor that improves readability but changes zero observable behavior is fundamentally different from one that touches actual logic, even if both look like similar-sized diffs. The prompt forces this distinction by asking for a Risk Score based on three specific factors: how much behavior the change touches, whether existing tests would catch a regression, and whether the change is purely structural or alters logic. This turns a vague sense of "this feels risky" into a repeatable judgment call the model has to justify.
Why low-risk fixes come first
Ordering suggestions from lowest to highest risk isn't just about safety — it's about momentum. Teams avoiding a bad function often avoid it entirely, including the parts that are trivially safe to fix. Banking a few low-risk wins first (renamed variables, deduplicated validation, replaced boolean flags with named parameters) builds confidence and shrinks the file before anyone has to make a harder call about the riskier structural changes.
Why it asks for test coverage before suggesting the risky stuff
The most common way refactoring goes wrong is skipping straight to the satisfying part — splitting a bloated function into clean, well-named pieces — without verifying the current tests actually exercise the code paths being changed. This prompt explicitly checks your stated test coverage and, if it's weak, tells you which test to write before attempting the higher-risk refactors rather than after something breaks in production.
Why the function split is flagged as its own PR
Splitting a multi-responsibility function into separate functions is usually the "real" fix everyone wants, but it's also the change most likely to alter subtle behavior — error handling paths, partial execution states, side-effect ordering. The prompt deliberately treats this as a separate, higher-risk refactor to be done only after the safer cleanups are merged and stable, rather than bundling it into the same change as the easy wins.
How to adapt it
Be specific about your constraints — a function called from 40 places behaves very differently under refactoring pressure than one called from a single test file. If you genuinely have zero test coverage, expect the model to recommend writing tests before touching anything past the first one or two low-risk items, and treat that recommendation as the actual answer rather than a formality to skip past.