A proof of concept usually gets scored against the requirements document, which means it gets scored against capability. Capability is the least informative thing a POC measures, because the vendor selected the environment and staffed the configuration.

The criteria worth scoring are the ones that survive contact with a real operations team. This scorecard weights them accordingly.

Run it rather than read it

The instrument below is also a browser-only tool that scores two columns side by side, flags operational vetoes, and exports to CSV: the POC Scorecard. Nothing you type leaves the page.

The rest of this piece is the reasoning behind the weights, which matters more than the arithmetic.

How to weight it

Three groups, weighted 60 / 25 / 15.

Operational behavior, 60%. How the tool acts when things go wrong, how much human attention it demands per week, and how it fits the stack you already run. This is where regret comes from.

Evidence and support, 25%. Whether the vendor can substantiate claims and whether support responded like a partner during the one period they were trying hardest.

Commercial and exit, 15%. What it costs to leave, and whether pricing survives your growth.

Capability appears nowhere as its own category. If the product cannot do the job, the evaluation ends before scoring.

The instrument

#CriterionGroupWeightWhat a 5 looks like
1Failure behaviorOperational8Degrades safely, alerts on its own outage, documented failure modes
2Alert qualityOperational8Signal per alert is high enough that the team reads them by week four
3Integration frictionOperational7Connected to your identity provider and SIEM without vendor engineering
4Weekly human costOperational7Under two hours per week of maintenance once tuned
5Performance at your volumeOperational6Tested at peak, not average, with headroom stated
6Upgrade behaviorOperational6At least one upgrade observed, or a public record of clean releases
7Multi-tenancy and scopingOperational6Permissions model matches your org structure without workarounds
8Observability of the tool itselfOperational6You can tell whether it is working without asking the vendor
9Claim substantiationEvidence7Every security claim on the website resolves to evidence you can read
10Documentation depthEvidence6Error semantics, limits, and migration paths are documented
11Support responsivenessEvidence6Real engineer, real answer, inside the stated window
12Reference comparabilityEvidence6Two references at your scale and stack, produced within a week
13Exit costCommercial8Data export in an open format, removal estimated under four weeks
14Pricing at 3x scaleCommercial7Cost curve modelled and acceptable at three times current volume

Score each criterion 1 to 5. Multiply by weight. The maximum is 470.

Scoring without fooling yourself

Two mechanics matter more than the criteria list, because both are where scorecards quietly become theatre.

Score independently, then compare. Have each participant score alone before any group discussion. Group scoring produces anchoring: the first number said out loud moves everything after it, and the loudest voice sets the tone. Independent scores also surface genuine disagreement, which is information. Two people scoring alert quality 2 and 5 have seen different things, and finding out why is worth more than the average.

Score the incumbent too. This is the discipline nobody applies and it changes outcomes. Run the same fourteen criteria against the tool you already have. A challenger scoring 390 looks compelling until the incumbent scores 375 and the migration costs a quarter. Without that baseline, every evaluation has a structural bias toward change, because the new thing gets measured and the current thing gets assumed.

A related trap: do not let the vendor see the weights before the POC. Weights shared in advance become a configuration target, and you will get a trial tuned to score well on exactly the eight dimensions you said mattered most.

How to read the output

The total is less useful than the shape.

Any criterion at 1 or 2 in the operational group is a veto candidate regardless of total. A tool that scores 420 with a 2 on alert quality will be ignored by the SOC within a quarter, and an ignored tool is a subscription, not a control.

Above 380 with no operational criterion under 3: buy it.

320 to 380: buy it if the gaps are in evidence and support, which improve with account attention, and do not if the gaps are operational, which do not.

Under 320: the evaluation should have ended earlier. Ask what kept it alive.

Worked example

Two products evaluated for the same job. Both cleared the capability bar, so neither appears in this table for what it can do.

CriterionWeightProduct AProduct B
Failure behavior84 (32)3 (24)
Alert quality83 (24)5 (40)
Integration friction75 (35)2 (14)
Weekly human cost73 (21)4 (28)
Performance at volume64 (24)4 (24)
Upgrade behavior64 (24)3 (18)
Multi-tenancy65 (30)3 (18)
Self-observability63 (18)4 (24)
Claim substantiation74 (28)5 (35)
Documentation depth65 (30)3 (18)
Support responsiveness63 (18)5 (30)
Reference comparability64 (24)2 (12)
Exit cost84 (32)2 (16)
Pricing at 3x scale73 (21)4 (28)
Total (max 470)361329

A twelve percent gap, which reads as a clear win for A until you look at shape.

Product B's 2 on integration friction and 2 on exit cost are both in the veto band. The integration score means somebody's quarter goes to connecting it. The exit score means the next renewal is a negotiation you have already lost. Product B is the more capable product on the dimensions people demo, and it is the more expensive product over three years.

Product A's weakest scores are alert quality at 3 and weekly human cost at 3, which are the two that most often decay after purchase. Those are worth a specific question to a reference customer at month eighteen before signing.

The total told you A. The shape told you what to negotiate, what to verify, and what will annoy your team in a year. That is the actual output of a scorecard, and reading only the total throws it away.

What a low score does not mean

A low score is a statement about fit with your environment at this moment, not about product quality. The most common cause of a poor operational score is that the tool assumes an organizational maturity you do not yet have: an on-call rotation to receive its alerts, an asset inventory to scope against, or an identity provider configured the way its connectors expect.

That is a real reason not to buy today. It is not a reason to conclude the product is bad, and it is worth writing the distinction into the evaluation record, because the same tool may be the right answer eighteen months from now and the next person should know why it was declined.

The scorecard also cannot see the two things that most often decide a security purchase after signature: whether your champion stays in the role, and whether the vendor gets acquired. Neither is scoreable. Both are worth a sentence in the risk section.