Better than what you have?
Plus, minus or same against a baseline. Coarse on purpose — and the real output is not the winner but the list of weaknesses to attack.
Better, same or worse than what you have now
Coarse on purpose. It asks only what people can answer reliably, and the ranking is a by-product of the real work.
Criteria
The matrix
How it scores
Each concept is rated against the baseline on every criterion: + better, − worse, blank the same. The score is the count of pluses minus minuses, or the weighted sum if weighting is on.
The comparison is always against the baseline, never between concepts. That is what keeps the judgements easy and consistent.
Coarseness is the feature
A Pugh matrix refuses to let you say "7 out of 10". That looks like a limitation and is the reason the method works.
People cannot reliably score seven concepts out of ten across eight criteria — the numbers drift, anchor on whatever was scored first, and change if you do it again next week. The same people can reliably say whether something is better than what they have now. Asking only what people can actually answer produces a more honest result than inviting precision nobody possesses.
The ranking is not the deliverable
The usual next step is not to pick the winner. It is to take the strongest concept and attack its minuses: can this option be changed so that its weaknesses go away, borrowing from the concepts that scored well there?
That produces a hybrid better than anything on the original list, and it is the entire reason the method exists. A Pugh matrix run once and filed has been used as a scoring sheet rather than a design tool.
Watch the baseline
Everything is relative to the baseline, so a weak baseline makes every concept look good and a strong one makes them all look marginal. Use the current approach rather than a strawman — the point is to find out whether change is worth it, and a rigged comparison answers a question you did not ask.