Selection Ends, Management Begins
A factory scorecard tracks four families — delivery (OTIF), quality (defect rate, AQL pass rate, claims), development (sampling speed and hit rate), and commercial behavior (communication, transparency, flexibility) — reviewed quarterly with volume allocation on the table.
The selection checklist answers 'should we start'; the scorecard answers 'should we continue, deepen, or leave' — and the second question recurs every quarter for the life of the program. The unmeasured relationship drifts on anecdote: the memorable late shipment outweighing nine on-time ones, the charming account manager covering a quality slide, the buyer's own busy season excusing patterns that would alarm them in a spreadsheet. Measurement corrects the anecdote in both directions — it defends good factories from one bad memory, and it exposes declining ones before the decline becomes a season.
The deeper function, which the best manufacturing partners understand and welcome: the scorecard is the relationship's shared truth. When both sides read the same OTIF percentage and the same defect rate, the conversation moves from blame to cause — the late order traced to the late approval, the defect spike traced to the material change nobody flagged. Programs that score their factories report the same paradox: the measured relationship is friendlier than the unmeasured one, because the numbers absorb the accusations that people otherwise trade. The factory that asks to be scored is telling you something about its confidence; the factory that resists is telling you something too.
The Four Scorecard Families
The architecture, kept deliberately at four families so the page stays readable: delivery (the promise kept — the flagship metric being OTIF, on-time-in-full: orders shipped on the confirmed date, complete, against the agreed 35-50 day window), quality (the standard kept — defect rate per inspection, first-pass AQL rate, warranty and claim incidence arriving from the field), development (the future built — sampling speed against the 6-10 day norm, sample hit rate, engineering contribution), and commercial behavior (the partnership kept — communication speed and honesty, transparency on problems, flexibility on the inevitable exceptions, pricing conduct at renewal).
Why four and not forty: every metric that survives on the card must drive a decision, and a scorecard with forty lines drives none — the review becomes a reading exercise instead of a decision meeting. The discipline is choosing the few numbers per family that carry the family's truth (the table below), instrumenting them so both sides can produce the data without argument, and letting the rest live in the operational detail where it belongs. The families, their core metrics and the question each answers:
| Family | Core metrics | The question it answers |
|---|---|---|
| Delivery | OTIF %; days late when late; partial-shipment rate | Can we plan seasons on their word? |
| Quality | Defect rate at inspection; first-pass AQL %; field claim rate | Does production match the approved standard? |
| Development | Sample lead time vs 6-10 day norm; first-sample hit rate | Is the sampling engine an asset or a bottleneck? |
| Commercial | Response time; problem-disclosure speed; flexibility record | Do they behave like a partner when it costs them? |
Delivery: the Promise Kept
OTIF is the scorecard's anchor because it is the metric everything else in the buyer's world hangs on: the buy plan, the freight booking, the retail floor-set and the event date all assume the confirmed ship date means something. Measured properly: on time (shipped on or before the confirmed date — with the confirmation itself documented, because the metric dies if the date can drift) and in full (the complete quantity — the partial shipment that forces a second freight bill is a delivery failure wearing a courtesy). The companion metrics that sharpen the reading: the distribution of lateness (three days late is a hiccup, three weeks is a season problem — the average hides which one you own), and the trend (one bad quarter is weather; three is a direction).
The honest sharing of the delivery ledger: late is not always the factory's number. The approval that sat in the buyer's inbox for nine days, the deposit that traveled late, the artwork that arrived after the cut date — the serious scorecard logs the buyer-caused delay separately, not as politeness but as accuracy, because the review that blames the factory for the buyer's calendar teaches the factory that the numbers are theater. The delivery section of the review is where both sides bring their timestamps — and the programs that run it honestly find the delivery number improving on both sides of the ledger at once.
Quality in Numbers
The quality family's three readings, taken at three altitudes: the inspection floor (defects found per AQL inspection — the rate and the first-pass percentage telling you whether the line holds standard without being forced), the loading dock (rejects, shortages, packing errors — the logistics-quality edge), and the field (warranty claims and defect-class returns arriving months later — the number that closes the loop between what the inspector saw and what the customer found). The three readings together describe the whole quality surface; any one alone lies — the inspection can pass while the field burns (a test gap), and the field can be quiet while the inspection burns (the rejects caught at cost, quality delivered by salvage rather than by process).
The numbers' proper use in the relationship: the trend over the absolute (every program's defect floor differs by product complexity — the feature-dense cart bag running a different baseline than the Sunday bag — so the card reads each factory against its own history, not against a universal bar), the failure-mode mix (the defects coded by type — stitching, color, component — so the review discusses the process, not the adjective), and the claim-resolution behavior (speed and fairness on the warranty claim being a quality metric too: the factory that resolves claims fast and fixes root cause is measurably different from the one that litigates every unit).
The Development Velocity
The family most scorecards forget, and the one that predicts the relationship's future: development is where next season comes from. The metrics: sample lead time (against the 6-10 day norm for a construction of known complexity — the trend telling you whether you are a priority or a queue position), first-sample hit rate (samples approved without a second round — the number that measures how well the factory reads a tech pack, and how well the buyer writes one), and engineering contribution (the factory's own suggestions that made the product better or cheaper — the pattern-maker's strap-angle fix, the material substitution that saved the colorway; counted, because what gets counted gets repeated).
Why velocity earns its place beside delivery and quality: a slow or sloppy sampling engine is a strategy tax — the line that misses the season window because the samples took five rounds, the trend that arrived at retail a year late because development consumed the margin of time. Measured over quarters, the development family also reads the relationship's health earlier than any other number: the factory deprioritizing your account shows up in sample speed two seasons before it shows up in production problems. The development line on the scorecard is the early-warning system disguised as an engineering metric.
The Commercial Behaviors
The soft family, scored as firmly as the hard ones: communication (response time on routine matters, and more tellingly, response time on problems — the factory that answers good news in an hour and bad news in a week has told you its conflict posture), disclosure (problems surfaced by the factory versus discovered by the buyer — the single most predictive commercial metric, because the partner who calls about the delay before you find it is pricing your trust as an asset), flexibility (the rush order accommodated, the mid-production change absorbed, the calendar rescue attempted — scored by disposition rather than by outcome, because the attempt reveals the posture), and pricing conduct (quotes that arrive explained, adjustments that arrive with evidence, MOQ and price conversations held in the open).
Scoring the unscorable, which is the family's craft: each behavior anchored to observable events rather than impressions (the disclosure metric counted from the actual sequence — who told whom, when — not from a vibe; the flexibility metric logged per request, granted or declined with reason), and the family's weight in the total kept honest (commercial behavior rarely outweighs delivery or quality in the math, but it is the family's trend that best predicts the hard numbers' future — the communication slide preceding the OTIF slide, the disclosure drought preceding the quality surprise). The soft numbers are the hard numbers' leading indicators; that is the entire justification for keeping them.
Weighting What Matters
The weighting decision, which is where the scorecard becomes a strategy document: the families do not weigh equally, and the weights change with the program's phase. The new relationship weights delivery and development (the unproven factory must first prove the calendar and the sampling engine — quality is gated by the inspection regime while volume is small), the mature relationship weights quality and commercial behavior (the proven calendar fades into assumption, and the long-run differentiators surface), and the crisis relationship — the season after a failure — weights the failed family until it re-earns its baseline.
The weighting table the review should open with, so the math is agreed before the numbers are read:
The weighting disciplines that keep the card honest: the weights set in advance (the quarter judged by the weights agreed before it began — the re-weighted scorecard being the anecdote creeping back in through the arithmetic), the gates kept separate from the weights (a catastrophic quarter in any single family flags the relationship regardless of the weighted total — the 92-score supplier with the safety failure is not a 92), and the weights themselves reviewed annually (the card serving the strategy, and the strategy moving).
| Program phase | Delivery | Quality | Development | Commercial |
|---|---|---|---|---|
| New relationship (year one) | 35% | 25% | 30% | 10% |
| Growth phase | 30% | 30% | 20% | 20% |
| Mature partnership | 25% | 35% | 15% | 25% |
| Recovery (post-failure) | Gate: the failed family must re-earn baseline | — | — | — |
The Quarterly Business Review
The meeting that turns the card into management: quarterly, ninety minutes, with the scorecard circulated a week ahead so the meeting discusses causes rather than reads numbers. The agenda that works: the card first (each family's trend, with both sides' data reconciled beforehand — the review is not the place to discover the two OTIF numbers disagree), the misses second (each significant miss with its root cause and its corrective action — owned, dated, and re-read at the next review), the plan third (next quarter's orders, the development calendar, the forecast shared as far ahead as it exists), and the relationship last (the commercial behaviors read aloud — the family most reviews skip and most relationships need).
The disciplines that keep the review from decaying into ceremony: the actions logged and re-opened (last quarter's corrective actions read first — the review that does not re-open its own commitments teaches both sides that commitments are ceremonial), the buyer's own misses on the same agenda (the late approvals and forecast misses logged beside the factory's — the symmetry being the review's entire moral authority), and the volume signal stated plainly (the allocation implication of the scores communicated, not implied — the factory told directly what its numbers are buying, because the scorecard without a volume consequence is a report card nobody studies for).
Scores Into Decisions
The translation layer where most scorecards die: numbers that never change a decision are expensive trivia. The decision rules, written in advance: the allocation band (scores mapping to volume share — the top band earning first-call on capacity and development slots, the middle band earning continuity, the bottom band triggering the improvement plan with dates), the development privileges (the high scorers seeing the new concepts first — the OEM pipeline flowing toward performance), and the commercial terms (the proven factory earning the smoother payment conversation — terms being a risk price, and the scorecard being the risk evidence).
The dual-sourcing judgment the card informs without dictating: a second source is insurance against concentration, and the scorecard prices the risk it insures (the consistently excellent single source with deep capacity earning continued concentration where the category allows; the volatile score line, the capacity ceiling or the strategic risk — the regional exposure — arguing for the second source regardless of any single quarter). The card's role is replacing the mood with the evidence: dual-sourcing decided in fear after one bad quarter is overcorrection; decided against three years of trend, it is policy.
When the Scorecard Says Leave
The reading nobody wants and the card exists to give early: the exit signals are patterns, not events — the two-family decline persisting past one improvement cycle (the improvement plan given its honest chance and failed), the disclosure drought (the problems now discovered rather than reported — the trust metric crossing the line from which relationships rarely return), and the development freeze (the sampling engine deprioritizing your account — the future being rationed before the present admits it). Any one pattern says plan; two together say move.
The leave executed as a discipline rather than a rupture: the transition run parallel (the replacement qualified through the full selection discipline before the incumbent is told — the bridge order covering the gap), the tools and property recovered per the molds-and-IP terms agreed at the start (the exit clause written in the honeymoon being the exit's entire ease), the final orders inspected harder, not softer (the last production of a departing relationship being quality's highest-risk window), and the professional close (the reasons stated once, factually, from the card — the industry being small and the next partnership beginning with this one's reputation for how it ends).
A Scorecard's First Year, Worked
The instrument's maiden year, from a composite program scoring a new stand bag partner: Q1 set the baseline (OTIF 92 percent — one shipment four days late on a lab-dip re-approval logged as buyer-caused; first-pass AQL 100 on two small orders; sampling at 8 days average with a 50 percent first-hit rate — the tech packs' fault as much as the factory's, and logged that way). Q2 showed the first real reading (OTIF 88 — a fabric mill delay disclosed nine days before the ship date, scored as a delivery miss and a disclosure win simultaneously, which is exactly how the two families are supposed to interact; the corrective action — earlier mill commitment — owned by the factory and re-read in Q3).
Q3 and Q4 told the year's story: the corrected delivery line (97, then 98), the quality floor holding (one inspection failure, root-caused to a new sewing team's first week, corrected before the next order — the trend line unbroken), the development engine warming (first-hit rate climbing to 75 as both sides learned the tech pack dialect), and the commercial family quietly compounding (the disclosure habit, the rush accommodation in November). The year-end review did what the card is for: the allocation moved the program's second chassis to this partner, the payment conversation softened on evidence, and the Q2 fabric action was confirmed closed. Four quarters of numbers bought what three years of dinners could not: certainty.
The Scorecard as a Relationship Asset
The closing frame, and the answer to the buyer who wonders whether scoring a partner is adversarial: the opposite, run properly. The scorecard protects the factory from the buyer's moods (the anecdote about the one late order dying against nine on-time quarters), it protects the buyer from the factory's charm (the account manager's warmth never again substituting for the defect trend), and it gives both sides the shared page that turns negotiation into engineering. The relationship that cannot survive measurement was never a partnership; the one that embraces it becomes measurably better — the numbers improving because they are read, which is the oldest management truth there is.
Which is why the best factories ask to be scored: they know their numbers, they know the numbers are good, and they know the buyer who measures stays longer, allocates rationally and renews on evidence rather than on fatigue. The scorecard is not the end of the handshake; it is how the handshake compounds — quarter over quarter, order over order, until the two organizations are running one production system with two names on the door.
Frequently Asked Questions
What is a factory scorecard?
A one-page ledger tracking a manufacturing partner's performance in four families — delivery (OTIF), quality (defect and AQL pass rates, field claims), development (sampling speed and hit rate), and commercial behavior (communication, disclosure, flexibility) — reviewed quarterly with volume allocation on the table.
What is OTIF and why does it matter?
On-time-in-full: orders shipped on or before the confirmed date, complete. It anchors the scorecard because your buy plan, freight booking, floor-set and event dates all assume the confirmed date means something. Track the lateness distribution and the trend, not just the average.
What metrics measure factory quality?
Three altitudes: inspection-floor results (defects found and first-pass rate at AQL inspection), loading-dock accuracy (rejects, shortages, packing errors), and the field loop (warranty claims and defect-class returns). Read trends against the factory's own baseline, coded by failure mode.
How often should I review factory performance?
Quarterly, ninety minutes, with the scorecard circulated a week ahead. Read last quarter's corrective actions first, discuss causes not numbers, share next quarter's forecast, and state the volume implication of the scores plainly — a scorecard without consequence is trivia.
Should the buyer’s own delays go on the scorecard?
Yes — separately logged, not blended. The approval that sat in your inbox for nine days is not the factory's OTIF miss. Symmetric honesty is the review's entire moral authority, and programs that run it find the delivery number improving on both sides at once.
How do I weight scorecard categories?
By program phase: new relationships weight delivery and development (calendar and sampling engine unproven); mature ones weight quality and commercial behavior. Set weights in advance, keep catastrophe gates separate from the weighted total, and review the weights annually.
What is a good first-sample hit rate?
It measures how well the factory reads your tech packs — and how well you write them. New relationships often start near 50 percent and climb toward 75-80 as both sides learn the dialect. Log the miss reasons; the trend teaches both parties.
When should I dual-source instead of scoring harder?
The scorecard prices the risk, it does not dictate the structure: volatile trends, a capacity ceiling, or strategic regional exposure argue for a second source regardless of any single quarter. Dual-sourcing decided on three years of trend is policy; decided on one bad quarter is overcorrection.
What are the exit signals on a factory scorecard?
Patterns, not events: two families declining past one improvement cycle, a disclosure drought (problems discovered, not reported), and a development freeze (your sampling deprioritized). One pattern says plan; two together say move — executed parallel, with molds recovered per contract.
Is it adversarial to score a manufacturing partner?
Run properly, the opposite: the card protects the factory from your moods and you from its charm, and turns negotiation into engineering. The best factories ask to be scored — they know their numbers, and they know measured buyers stay longer and allocate rationally.
What is a quarterly business review with a factory?
The meeting that converts scores into management: the card's trends, each miss with root cause and dated corrective action, next quarter's orders and development calendar, and the commercial behaviors read aloud. Actions are logged and re-opened — ceremony is the decay mode.
How do scores translate into commercial terms?
Write the decision rules in advance: allocation bands mapping scores to volume share, development privileges (top scorers see new concepts first), and payment-term softening on evidence — terms are a risk price, and the scorecard is the risk evidence.