What a scorecard is for, and what it is not
Three activities get confused with one another, and running them as a single process makes all three worse.
A supplier can pass a rigorous audit, clear PPAP on every part, and still deliver badly for two years. The audit looks at the system; the scorecard looks at the output. You need both, and the scorecard is the one that runs continuously.
The purpose is not to produce a report. It is to make three decisions defensible: where new business goes, which supplier gets development attention, and when a supplier is removed. If your scorecard does not feed those three decisions, it is administrative work generating a number nobody acts on.
A full scorecard is justified for strategic suppliers carrying meaningful spend or single-sourced critical parts. For everyone else, on-time delivery and lot acceptance on a simple annual review is sufficient. Extending the full apparatus across thirty suppliers guarantees it is done poorly on the four that matter.
The five metrics that matter for magnets
Scorecards bloat. Every function wants its concern represented, and the composite ends up averaging twelve weakly-correlated numbers into something with no signal. Five is close to the practical maximum, and for magnets these five cover the ground.
| Metric | What it measures | Source | Why it matters for magnets specifically |
|---|---|---|---|
| On-time delivery | Receipts inside the agreed window | Your ERP receipt dates | Lead-time variability drives most of your safety stock; this is the metric that moves working capital |
| Lot acceptance | Share of lots accepted without deviation | Incoming inspection records | Magnets fail in distinct modes — magnetic, dimensional, coating — and lot-level acceptance captures all three |
| Documentation | Certifications complete, correct and on time | Receiving and quality records | A wrong or missing certificate of conformance stops a regulated build as effectively as a wrong part |
| Responsiveness | Quote turnaround, RFQ hit rate, issue closure time | Purchasing and quality logs | The leading indicator; responsiveness degrades before delivery does |
| Commercial | Price against benchmark, index adherence, cost initiatives | Should-cost model and index clause | Keeps the scorecard from rewarding a supplier who is merely expensive and reliable |
Responsiveness is the one most often left out and the one with the most predictive value. Quote turnaround stretching from three days to two weeks, or an unanswered corrective action request, almost always precedes a delivery problem by a quarter or more. It is soft to measure and worth measuring anyway.
Every metric needs an owner and a system of record identified before the scorecard launches. A metric whose data is assembled by hand each quarter will be assembled late, then estimated, then quietly dropped. If the number cannot be pulled from a system, either fix the system or remove the metric.
Defining on-time delivery so it cannot be gamed
On-time delivery is the most-quoted supplier metric and the most inconsistently defined. Two companies scoring the same supplier on the same shipments routinely produce numbers twenty points apart. Four decisions determine the answer, and all four should be written into the supply agreement rather than left to whoever built the report.
The date-basis choice matters most. Scoring against the promise date measures whether the supplier keeps commitments. Scoring against your request date measures whether they meet your needs. A supplier can run 100% against promise dates while quoting sixteen weeks on a part you need in eight, and a scorecard showing perfect delivery on a part that is chronically late is a scorecard measuring the wrong thing.
The defensible answer is to track both, and to treat a widening gap between them as its own signal. Score the composite on promise-date adherence, because that is what the supplier controls; report request-date adherence alongside it, because that is what your plant experiences.
Suppliers under delivery pressure ship early, which improves their number while transferring carrying cost, storage and obsolescence risk to you. If your window has no early bound, you are paying for the metric you asked for. On magnets there is an additional consideration: early arrival of magnetized stock consumes segregated storage space that is usually already tight.
Measuring magnet quality properly
Generic parts-per-million defect rates travel badly to magnets. Lot sizes are large, sampling is statistical rather than exhaustive, and the failure modes are not equivalent to each other. A PPM figure averages a cosmetic coating blemish and a grade shortfall into one number, and those two findings have nothing in common.
Score lot acceptance, not PPM
Lot acceptance — the share of received lots accepted without deviation, rework or waiver — maps directly onto how magnets are actually inspected and onto the decision your quality team makes. Track PPM underneath it for parts under statistical control, but let the composite run on lots.
Weight the defect classes differently
| Class | Typical findings | Detectability at incoming | Weight |
|---|---|---|---|
| Magnetic | Grade below spec, low remanence, coercivity shortfall, wrong magnetization direction | Requires a helmholtz or hysteresisgraph; frequently escapes basic inspection | Highest |
| Dimensional | Out-of-tolerance thickness, flatness, perpendicularity, chamfer | High — caught by routine gauging | High |
| Coating | Adhesion failure, thickness shortfall, corrosion in salt spray, edge coverage | Moderate; adhesion and corrosion need destructive or timed tests | High |
| Documentation | Missing or wrong CoC, absent test data, unstated material substitution | High, if anyone actually reads the certificate | Moderate to high |
| Packaging | Inadequate separation, damage in transit, incorrect labelling | High | Lower, unless it caused damage |
Magnetic non-conformance ranks highest because it is the one most likely to reach your assembly line undetected and to fail in the field rather than at receiving. A supplier whose defects are all dimensional has a process control problem; a supplier whose defects are magnetic has a material or grade problem, and those are different conversations. The testing guide covers what each measurement actually requires, and incoming inspection and acceptance covers building the receiving plan that generates this data.
A grade or coating substitution that performs adequately but was never disclosed should be scored as a serious documentation non-conformance even when the parts work. It indicates the supplier is making unilateral changes to a controlled specification, and the next substitution may not be benign. This is the single most useful thing a magnet scorecard can catch that an audit will not.
Weighting and scoring the composite
Weights encode what you care about, and suppliers will optimise against them precisely. Publish them. A supplier who knows delivery carries thirty-five points and documentation fifteen will allocate their attention accordingly, which is the entire point.
| Metric | Weight | 100 points | 50 points | 0 points |
|---|---|---|---|---|
| On-time delivery | 30% | ≥ 98% | 92% | ≤ 85% |
| Lot acceptance | 30% | 100% | 96% | ≤ 92% |
| Documentation | 15% | No findings | 2 findings | ≥ 5 findings |
| Responsiveness | 15% | Quotes ≤ 3 days, CAs closed on time | Quotes ≤ 7 days | Quotes > 10 days or CAs overdue |
| Commercial | 10% | At or below benchmark, index adherence clean | Within 5% of benchmark | > 10% above benchmark |
Interpolate linearly between the anchors and cap at 100 so that outstanding performance on one metric cannot mask failure on another. The scoring bands should be steep enough to discriminate: if every supplier lands between 88 and 94, the anchors are set too generously and the scorecard is not doing any work.
Use a rolling window
Score on a rolling twelve months, recomputed quarterly. A single quarter contains too few magnet deliveries to be statistically meaningful — on a part shipping monthly, one late delivery moves quarterly on-time performance by 33 points. Rolling twelve months damps that noise while still responding to a real trend within two quarters.
Reserve the right to place a supplier in the red band regardless of composite score for a single serious event: a safety-relevant escape, a falsified certificate, or an undisclosed change to a controlled process. Averaging those away is the fastest route to a scorecard nobody trusts.
Review cadence and what happens in each band
The score is an input to a conversation. The conversation needs a schedule and a defined outcome per band, agreed in advance, or the review becomes a presentation of numbers followed by nothing.
| Band | Score | Cadence | Required action |
|---|---|---|---|
| Green | 90–100 | Annual business review | Eligible for new business and longer agreements; discuss cost and capacity roadmap |
| Amber | 80–89 | Quarterly review | Written improvement plan against the specific metric, with dates and an owner |
| Red | < 80 | Monthly until recovered | Formal corrective action, no new business, second source activated in parallel |
| Probation | Red for two consecutive quarters | Monthly, escalated | Exit plan drafted and requalification of the alternative funded |
What belongs in the quarterly review
That final row is the one that separates a review from an interrogation. Late drawings, unrealistic request dates, and a frozen window shorter than the supplier's material lead time are your contributions to their delivery number, and a supplier who cannot raise them will manage the metric rather than fix the problem.
Making the score mean something
The reason most scorecards fail is not measurement. It is that a supplier can score 68 for four consecutive quarters and observe that their volume did not change. Once that has been demonstrated once, the scorecard is understood correctly by everyone as a reporting exercise.
| Consequence | Trigger | Practical constraint |
|---|---|---|
| New business allocation | Green band, current qualification | The most effective lever and the easiest to apply; costs nothing to implement |
| Volume shift between qualified sources | Amber for two quarters | Only works if the alternative is genuinely qualified on that part number |
| Longer agreement or index terms | Sustained green | Valuable to the supplier, cheap for you, and it rewards the right behaviour |
| Development support | Amber with a credible plan | Reserve for suppliers you intend to keep; it is an investment, not a courtesy |
| Removal from approved list | Probation without recovery | Genuinely expensive on magnets — tooling, requalification and lead time all reset |
The honest limits
Magnet supply constrains what a scorecard can enforce. Tooling is part-specific, so shifting volume requires either a duplicate die or a full requalification at the alternative. Grades and coatings are not perfectly interchangeable between producers even at the same nominal specification. And where a part is genuinely single-sourced, a red score buys you a corrective action and very little leverage beyond it.
Which is why scorecard design and sourcing strategy are the same exercise. If you want consequences available, you need a qualified alternative on the parts that matter — funded and requalified before the score turns red, not after. The second-source qualification guide covers what that costs and how long it takes; the honest planning assumption is several months and a tooling charge, which is exactly why it has to be started while the incumbent is still performing.
A distributor holding domestic inventory and a factory running your tooling are not comparable on the same metrics. Lead time and MOQ mean different things in each case, and a stocking distributor scored on production metrics will look artificially good while a factory scored on availability will look artificially bad. Segment the scorecard by supplier role — the category strategy guide sets out those roles.
