The RFP scoring matrix, so the best deck doesn’t beat the best plan
An RFP scoring matrix is a weighted scorecard for comparing vendor proposals. Each criterion gets a weight reflecting how much it predicts a good outcome, each proposal gets a 1-to-5 score per criterion, and the weighted totals make the comparison explicit. Its job is to stop the best-looking proposal from beating the best plan, which is what happens when a committee compares documents by feel.
Download the Matrix↓Word + PDF, editable weights, no email required
Why unweighted comparison fails
Put three proposals in front of a committee with no scoring structure and two things reliably happen. First, the comparison drifts to whatever is easiest to compare, which is always the price and the page count, never the content plan. Second, the most polished document sets the reference point, and everything else gets read against it. Neither has much to do with which firm will still be doing good work for you in year two.
Weighting fixes the first problem by making the hard-to-score criteria count for more than the easy ones. Scoring before discussing fixes the second: each evaluator commits to numbers independently, and the divergences become the agenda for the meeting instead of the loudest voice.
One rule matters more than the rest: set the weights before proposals arrive, and write them into the RFP itself so the agencies know the game is fair. Weights adjusted after reading the proposals aren't weights, they're a justification.
The criteria, and why these weights
Eight criteria, prefilled to total 100. Edit them to your project; the prefilled version reflects what we’ve seen predict good outcomes on B2B website projects, and price is deliberately last.
| Criterion | Why it’s weighted this way |
|---|---|
| Understood the actual problem (20%) | The single best predictor. A proposal that names what's costing you buyers, with evidence from your own site, was written for you. One that could've been sent to anyone will be executed the same way. |
| Relevant work with outcomes (15%) | One comparable project with what changed after launch beats a large portfolio. You're scoring evidence the firm has solved your shape of problem, not their taste. |
| Content plan (15%) | The most common reason projects stall. Score who writes, who interviews your team, and who owns the deadline. 'Client to provide copy' scores a 1. |
| SEO migration plan (15%) | If the current site ranks, this is where a redesign quietly destroys value. Score whether the proposal names the URL inventory, redirect map, baseline, and monitoring without being asked. |
| Team and continuity (10%) | Score whether you know who actually does the work, and what happens if they leave. The people in the pitch are rarely the people in the project. |
| Timeline realism (10%) | Milestones that account for your team's availability, especially around content review. A schedule that assumes instant approvals is a schedule that slips. |
| Ownership (10%) | Code, content, domain, and accounts in your name at the end. Anything else is a dependency priced into year two. |
| Price completeness (5%) | Not the price: the completeness. Score whether exclusions are stated and priced. The number itself already influences everyone enough without extra weight. |
The scoring matrix
One page, three proposal columns, weights prefilled and editable. Score 1 to 5, multiply by the weight, total the column.

Use the Word version: the weights are meant to be edited, and a matrix you didn’t adjust to your project is a form, not a decision.
A worked example
Three fictional proposals for the same B2B redesign. Proposal A is $28,000 from a firm that quoted fast: polished deck, big portfolio, no mention of your current site. It scores 2 on problem understanding, 2 on the content plan (“client to provide copy”), 1 on SEO migration. Weighted total: around 250 of 500.
Proposal B is $61,000: strong on everything, scores 4s and 5s across the board, weighted total around 440. Proposal C is $39,000: opens by quoting your own homepage back at you and naming what it fails to say, scores 5 on problem understanding, 4 on content because they write it, 4 on migration, 3 on team continuity because it’s a small firm. Weighted total: around 420.
The matrix doesn’t make the call between B and C; a twenty-point gap on 500 is a tie. What it did is eliminate A before the price anchored anyone, and turn the real decision into a named trade: B’s depth against C’s attention, at a $22,000 difference. That’s a conversation a committee can actually have.
Running the scoring meeting
Score before you meet, not during. Each evaluator fills in their own copy independently, because a shared screen turns scoring into anchoring: the first number said out loud becomes the number. Then compare columns, and spend the meeting only on the rows where scores diverge by two or more. Those divergences are the entire value of the exercise: they surface what one reader caught and another missed, and they do it before the contract is signed rather than in month three.
Keep the completed matrices. If the decision is ever questioned, by a board, an owner, or the team living with the site two years later, a dated scorecard with named evaluators is the difference between a defensible process and a recollection.
When the cheapest proposal scores highest
Sometimes it should. A firm that understood the problem, writes the content, and handles the migration can genuinely be the cheapest, especially against competitors padding scope. Before you celebrate, check the exclusions row: the common reason a low quote wins on paper is that content, redirects, and post-launch support are outside it, and each returns as an invoice with worse leverage than you have today. If it still wins with the exclusions priced in, take it.
And if the scores cluster, break the tie with the two things the matrix can't capture: the reference calls, and which team asked you the better questions. For what to listen for on those calls, the agency selection guide covers the questions and what good answers sound like.
Common questions
What is an RFP scoring matrix?
A weighted scorecard for comparing vendor proposals. Each criterion carries a weight that reflects how much it predicts a good outcome, each proposal is scored 1 to 5 per criterion, and the weighted totals put the comparison on paper. It exists to stop the best-presented proposal from beating the best plan.
How do you score an RFP?
Set the criteria and weights before proposals arrive, have each evaluator score independently, then compare notes where scores diverge. Score 1 to 5 per criterion, multiply by the weight, and total. The divergences are the useful part: they surface what one reader caught and another missed.
What criteria should be in a vendor evaluation for a website project?
The ones that predict the outcome: whether the vendor understood your specific problem, comparable work with results, who writes the content, the SEO migration plan, team continuity, timeline realism, what you own at the end, and the completeness of the price. Our prefilled weights put problem-understanding highest at 20% and price lowest at 5%.
Why weight the criteria instead of scoring them equally?
Because unweighted scoring quietly becomes price scoring. When everything counts the same, the criteria that are easy to score (price, page count) dominate the ones that are hard to score and matter more (problem understanding, content plan). Weights force the decision you'd defend later.
What if the cheapest proposal scores highest?
Then it might genuinely be the best, but check one thing first: the exclusions. The most common reason a low quote wins on paper is that content, redirects, and post-launch support are not in it, and each comes back later as an invoice. If it still wins with the exclusions priced in, take it with a clear conscience.
How many people should score each proposal?
Two or three, independently, then reconcile. One scorer inherits their own blind spots; five turns scoring into a meeting. The person who talks to buyers and the person who will run the site day to day catch different problems, so make sure both are represented.