Supplier evaluation criteria that actually predict performance
Ask a buying team what their evaluation criteria are and you will usually hear the same three words: price, quality, delivery. Ask how they scored the last award against those criteria and the room goes quiet. The criteria existed, but the decision was made the usual way, someone read the bids, formed a view, and then wrote the scores to match it.
That is the quiet failure of most supplier evaluation. The criteria are not wrong, they are just decorative. They are written after the bids arrive, applied unevenly, and bent toward the supplier someone already prefers. The fix is not a longer list. It is a shorter list of criteria that actually predict what happens after the award, applied the same way to every bid.
Prediction, not justification
A useful evaluation criterion answers one question: does this tell me something about how the supplier will perform after I award? That test kills most of the filler that accumulates in evaluation templates.
"Quality of proposal document" fails the test. A polished PDF predicts a good bid writer, not a good supplier. "Years in business" mostly fails it too; longevity is weak evidence compared to what a supplier has actually delivered recently. Meanwhile the criteria that do predict performance are often the ones nobody scores formally: whether the supplier answered clarification questions quickly and precisely, whether their price is explainable, whether buyers who worked with them would do so again.
So before writing a scorecard, run every candidate criterion through the prediction test. If it would not change your confidence about delivery, it does not belong on the list, however traditional it looks.
Gates first: the criteria you should never score
Some requirements are not more-or-less, they are yes-or-no. A required certification. A delivery date your operation cannot survive missing. A payment term your finance team will not sign. These belong in a different category from scored criteria, because averaging hides them.
Here is the failure mode averaging creates. A bid misses a mandatory certification but is 30 percent cheaper than everyone else. On a weighted scorecard, the price score is so strong that the bid still ranks first, and now you are arguing with your own scorecard about why you cannot award to it. The scorecard has become the enemy of the requirement.
Gates fix this. A bid that fails a gate is out, and its price never enters the comparison. On VEXORS these live in the request as show-stopper requirements, and a bid that fails one is disqualified regardless of its number. Decide your gates when you write the request, not when a tempting bid forces the question. We covered this mechanism in more depth in how AI bid scoring works.
The criteria that earn their place
After the gates, most purchases need five or six scored criteria, not fifteen. These are the ones that keep proving themselves:
| Criterion | What it predicts | What to look at |
|---|---|---|
| Total cost | What you will actually pay | Landed cost including delivery, installation, and payment terms, not the headline unit price |
| Specification compliance | Whether you get what you asked for | Line-by-line match against your bill of quantities, deviations flagged and explained |
| Delivery capability | Whether it arrives when promised | Lead time evidence, capacity for your volume, named logistics arrangements |
| Track record | How similar work went before | Completed contracts of similar scope, not marquee logos from a different category |
| Peer signal | Whether past buyers would return | Ratings tied to completed contracts, patterns across many jobs rather than one curated reference |
| Responsiveness | What the working relationship will feel like | Speed and precision of answers during clarification, before any contract exists |
Two of these deserve a note. Total cost, not unit price, because the cheap unit price with expensive delivery, long payment terms, and a costly installation is one of the oldest tricks in the book, and a criterion named "price" invites it. And responsiveness, because the supplier's behaviour during the sourcing process is a free preview of the relationship. A bidder who takes a week to answer a simple clarification while competing for the work will not get faster after winning it.
Company-level checks, verification, registration, credentials, sit underneath all of this and are covered in our guide to vetting suppliers beyond price. The scorecard assumes the companies in it deserve to be there.
Weights: decided before, not after
Weights are where evaluations quietly get bent. Choose them after reading the bids and they will drift toward whichever supplier impressed you first; that is not dishonesty, it is just how anchoring works. The protection is mechanical: set the weights when you write the request, record them, and leave them alone.
How to choose them honestly? Ask what failure costs you.
- Commodity purchase, many substitutes, easy to switch: price can carry 50 or 60 percent. If a delivery slips you buy elsewhere and lose little.
- Operationally critical, hard to substitute, failure stops your line: delivery capability and track record together should outweigh price. Saving 8 percent means nothing against a week of standstill.
- Long-term service relationship: responsiveness and peer signal move up, because you are buying years of working together, not a shipment.
There is no universal split, and that is the point. The weights are a statement about what this purchase can and cannot tolerate, made while your judgment is still clean.
A scorecard you can copy
For a mid-sized RFQ with a bill of quantities, a defensible starting structure looks like this:
Gates (pass or fail): valid trade licence and any mandatory certification, delivery to site by the required date, acceptance of your core commercial terms.
Scored criteria:
| Criterion | Weight |
|---|---|
| Total cost against the bill of quantities | 40% |
| Specification compliance and flagged deviations | 20% |
| Delivery capability and lead time evidence | 15% |
| Track record on similar completed work | 15% |
| Responsiveness during the process | 10% |
Adjust the numbers to your risk, but keep the shape: gates separate from scores, total cost rather than unit price, and no criterion so small it exists only for show. Five criteria scored honestly beat twelve criteria scored from fatigue.
Consistency is the hard part, and the automatable part
A good scorecard applied unevenly is worse than no scorecard, because it produces confident numbers that mean different things for different bids. And unevenness is the natural state: the eighth bid on a long afternoon is not read the way the first one was.
This is exactly the part that structure and automation fix. On VEXORS, the request carries your gates as show-stoppers and your questions as a questionnaire every bidder answers in the same shape, bids come back line by line against the same bill of quantities, and AI scoring reads every qualifying bid against the criteria you set, identically, whether it is bid one or bid thirty. You still make the award, and you can still override the ranking for reasons the data cannot see. What you stop doing is pretending a tired human reads bid eight the way they read bid one.
The short version
Write criteria that predict performance, not criteria that decorate a decision. Put the non-negotiables in gates so no price can drag a non-compliant bid back into the running. Score total cost, compliance, delivery evidence, track record, peer signal, and responsiveness, weighted by what failure would actually cost you, with the weights fixed before any bid arrives. Then apply the whole thing evenly, which is precisely the step worth handing to a machine.
Ready to run an evaluation your scorecard can defend? See how VEXORS structures requests, bids, and scoring so the criteria you set are the criteria that decide.
Frequently asked questions
- What criteria should I use to evaluate supplier bids?
- Start with pass-or-fail gates: specification compliance, mandatory certifications, and delivery feasibility. Then score what remains on total cost rather than unit price, evidence of delivery capability, track record on similar work, peer ratings from past buyers, and responsiveness during the process itself. The test for any criterion is whether it predicts performance after award, not whether it is easy to score.
- How should I weight price against quality and delivery?
- It depends on what failure costs you. For a commodity purchase with many substitutes, price can carry half the weight or more. For anything where a delivery failure stops your operation, delivery evidence and track record should outweigh price. The honest method is to set the weights before any bid arrives, based on what would actually hurt, and then leave them alone.
- When should a criterion be pass-or-fail instead of scored?
- When you would never award to a supplier who misses it, no matter how good the rest of the bid is. A required certification, a hard delivery date, a mandatory commercial term. Making these gates rather than weighted scores stops a very low price from dragging a non-compliant bid back into contention.
- Should I tell suppliers what the evaluation criteria are?
- Share the criteria and any pass-or-fail requirements, because suppliers who know what matters write better, more comparable bids. Exact weights are your choice: sharing them increases transparency, while keeping them private discourages bids engineered to the formula. Many teams share the criteria list and keep the weights internal.
Run procurement the modern way
Create a free account and see how structured sourcing and supplier trust come together on VEXORS.
Get started free