The instrument weights operational behavior at 60% of the total, evidence and support at 25%, and commercial and exit terms at 15%. Capability is not scored at all: if the product cannot do the job, the evaluation ends before scoring starts, and capability is the dimension a vendor controls most tightly during a trial.

Why these weights. The failures that make security teams regret a purchase are cumulative and operational: alert fatigue at month four, rule maintenance at month eight, an upgrade that breaks an integration at month eleven. A thirty-day trial cannot observe any of those directly, but it can observe their leading indicators, which is what the operational group measures.

How to read the output. The total is less useful than the shape. Any operational criterion scoring 1 or 2 is a veto candidate regardless of the total, because a tool the team stops reading is a subscription rather than a control. Above 380 with no operational criterion under 3, buy it. Between 320 and 380, buy it if the gaps sit in evidence and support, which improve with account attention, and do not if they sit in operations, which do not.

Score the incumbent too. The tool lets you run two columns for this reason. Without a baseline, every evaluation carries a structural bias toward change, because the new thing gets measured and the current thing gets assumed.

What it cannot see. Whether your champion stays in the role, and whether the vendor gets acquired. Neither is scoreable and both decide outcomes after signature.