I built this vendor scorecard template for a real evaluation earlier this year. Seven pass/fail gate items, fourteen weighted criteria, a clean grid, one row per vendor. Then I sat with the team as they worked through it.
They walked into the room and wrote verdicts straight into the decision row. Zero gate items answered. Zero scores entered. The instrument was sitting open on the screen and it did not touch the conversation once.
That failure is the most useful thing I know about vendor scorecards, and I will come back to it at the end, because the fix is not a better template.
The template still matters though, and the one most people use is broken in a specific way. A single weighted average lets a vendor average out a flaw you could never live with. Score high enough on the things a demo shows well and you can fail the one thing that ends the relationship in month four, and still finish with a number that looks fine on a slide.
So this is a two-stage instrument. Stage one is a pass/fail gate of non-negotiables, and any single Fail ends the evaluation there, no score calculated. Stage two is weighted 0 to 5 scoring, survivors only. Both stages are on this page in full, as tables you can copy. No form, no download, no email wall.
It is written for whoever actually owns the consequences of the choice: the IT manager or CIO buying something that will sit inside the estate, and the service provider buying something they will resell and stand behind. The instrument is the same for both. Two criteria change wording depending on which you are, and I flag them where they appear. I spent a decade evaluating companies in investment banking and private equity before I started running growth for IT services and cyber firms, and the gate came straight off the private equity side. Diligence has always worked this way, as a screen before a score. Software and supplier scorecards mostly forgot.
Why most vendor scorecards fail
The standard vendor evaluation scorecard is a list of criteria, a weight per criterion, a score per vendor, and a weighted average at the bottom. That is a reasonable instrument for choosing between options that are all acceptable, and a terrible one for finding out whether an option is acceptable at all. Averaging is designed to let strength compensate for weakness. That is the whole point of an average, and it is also the failure.
Run the arithmetic. Take the fourteen weighted criteria below, weights summing to 32, five points per criterion, so 160 points on the table. Now score a vendor that demos beautifully.
Automation depth 5. Scope 5. Price at your volume 4, cost fit 4, billing fit 4, integration 4, support 4, prevention alignment 4, management overhead 4, onboarding 4. Price scalability at 2x volume 2, because nobody could answer that one. Data location 2. Security posture and certifications 1, because there is a SOC 2 Type II report and nothing behind the automation claims. Peer and reference signal 0, because there are no reference customers at your scale and no independent reviews anywhere.
That vendor scores 114 out of 160. Seventy-one percent. It goes on the slide as a strong contender, and it has zero verifiable customers and no third-party validation of the capability you are buying. The two marks that should have ended it, a zero and a one, got absorbed by eight fours and two fives, most of them earned in a controlled demo.
That is not a scoring error. The spreadsheet did what a weighted average does. The mistake was asking a compensating instrument to do a disqualifying job.
Some flaws are not weights. No exit terms in writing is a hostage situation you signed up for, not a low score on flexibility. An unverifiable AI claim is the whole product being unproven, not a small deduction on capability. Those belong in a gate, and a gate has exactly two outputs.
Stage one: the gate
Seven items. Each one is Pass or Fail. One Fail and the evaluation is over, and nobody scores anything. The point of writing them down in advance is that it is hard to argue yourself past a Fail you defined before you met the salesperson.
If you have run a vendor risk assessment before, this will look familiar, and that is deliberate. The difference is where it sits. Risk assessment usually happens after selection, as a compliance step on a decision already made, which is how a vendor due diligence checklist turns into paperwork. Run the same questions first and they stop being paperwork and start being a filter.
| # | Gate item | Pass condition | How to verify |
|---|---|---|---|
| 1 | Independent attestation of automation and AI claims | SOC 2 Type II plus a technical report covering the detection and triage claims specifically | Ask for the report under NDA and read the scope section. Marketing material, a trust page, and a benchmark the vendor ran on itself are all Fail |
| 2 | Reference customers at your scale | Two to three customers running it in production at roughly your size and complexity | You call them yourself. Not a vendor-run reference call, not a logo wall, not a case study with an unnamed customer |
| 3 | No selling around you, in writing | If you resell: no direct approaches or quotes to your clients. If you buy in-house: no selling directly into your business units behind IT | A clause in the agreement, not a verbal assurance from the rep. Ask what their sales team is compensated on, because that predicts the behaviour better than the clause does |
| 4 | Integration tested in a sandbox before signature | One alert creates one ticket. No duplicate-incident loops. No second portal your team has to live in | Run it in a sandbox against your real ticketing system before you sign, and watch what lands |
| 5 | Exit terms in writing | Notice period, data export format, offboarding cost, and deletion of your data all specified | Read the termination clause. If export format is unspecified, that is a Fail, not a detail for later |
| 6 | Liability and subprocessor terms in writing | Who at the vendor can reach your customers' systems, from where, with what audit trail, plus the disclosure language you can put in your own client agreements | Subprocessor list, access model, and the exact wording you are allowed to pass downstream |
| 7 | Contract term of 12 months or less | 12 months or shorter preferred. A required 36-month lock-in fails unless the pricing is exceptional | The order form, not the sales conversation |
Two of these do more work than the other five put together.
The attestation question. Ask for independent validation of the specific claim they are selling, and the answer sorts the market in about ninety seconds. Almost every security vendor has SOC 2 Type II now, and SOC 2 Type II says the company runs decent internal controls. It says nothing about whether the AI triages correctly. If a vendor tells you its agents resolve 95 percent of Tier-1 cases, the question is who measured that other than the vendor. Usually the honest answer is nobody, and they will say so if you ask directly enough. Not a scandal in a young category. Just a Fail on a gate you wrote before you walked in.
The exit question. Ask how you get your data out, in what format, at what cost, and how long deletion takes. This one is diagnostic beyond its own answer. A vendor with a rehearsed answer has been through offboardings and has customers mature enough to demand it. A vendor that has never been asked will improvise, and you will hear the improvisation. Ask it on the first commercial call, before the sales team is emotionally invested in your logo.
Stage two: the weighted score
Only vendors that pass all seven get here. Fourteen criteria in three blocks, weights from 1 to 3, weights summing to 32, so 160 points available.
| Block | Criterion | Weight | What a 5 looks like |
|---|---|---|---|
| Price and commercials | Price at our volume | 3 | Quoted at our actual seat or ingestion count, not list, and it clears our budget |
| Price and commercials | Cost fit, or margin fit if you resell | 3 | In-house: it fits the budget line it has to come out of, all in, including the internal time to run it. Reselling: we can resell at our standard margin without repricing our own service |
| Price and commercials | Billing fit | 2 | Matches how we bill clients: same cycle, same unit, no true-ups we cannot pass on |
| Price and commercials | Price scalability at 2x volume | 2 | We have the price at double today's volume in writing, and it does not step up |
| Reputation, compliance and fit | Peer and reference signal | 2 | Multiple independent operators at our scale using it in production and willing to say so |
| Reputation, compliance and fit | Security posture and certifications | 2 | Current certifications, published subprocessor list, clean recent incident history |
| Reputation, compliance and fit | Support quality and responsiveness | 2 | Named escalation path, tested response times, humans who know our environment |
| Reputation, compliance and fit | Data location and sovereignty | 1 | Data resides where our client contracts require, documented, with no silent regional failover |
| Reputation, compliance and fit | Integration with the existing stack | 3 | Live, generally available integrations with the tools we actually run, not roadmap entries |
| Reputation, compliance and fit | Scope, meaning what they take off your plate | 3 | Replaces a defined set of work we do today, and we can name the hours |
| Operational | Automation depth | 3 | Automates end to end on the cases that matter, with the decision trail visible to us |
| Operational | Prevention alignment | 2 | The vendor tunes to reduce alert volume over time rather than earning more from noise |
| Operational | Management overhead added | 2 | No new full-time babysitting, no second console, no separate on-call rota |
| Operational | Onboarding, migration and training | 2 | Fixed-price onboarding with a dated plan and named people |
Scoring 0 to 5 honestly is the part teams get wrong, so here is the scale I use.
0 means fails or unknown. Not "probably fine". Not "they said they would send it". If you do not have the answer in hand it is a zero, and the zero is information: it tells you how much of this decision you are making blind. Teams that score unknowns as 3 turn a scorecard into a mood ring.
1 is materially worse than what you run today. 2 is workable with a named compensating control, and write that control next to the score or you will forget it existed by the time you are live.
3 is genuinely fine, and most criteria for most decent vendors are 3s. A scorecard where everything is a 4 or 5 got filled in after the decision.
4 beats your current state in a way you can measure. 5 is a reason to buy on its own. Hand out very few 5s. If a vendor has four of them, check whether you scored the demo or the product.
Score every vendor in one sitting with the same people. Scores drift when one gets scored the day after its demo and the competitor gets scored three weeks later from memory.
The full vendor scorecard template
Both stages, one block, copy it straight out.
Stage one: gate. Any Fail ends the evaluation.
| # | Gate item | Verdict | Owner | Date answered | Evidence |
|---|---|---|---|---|---|
| 1 | Independent attestation of automation and AI claims (SOC 2 Type II plus a technical report covering the detection and triage claims) | Pass / Fail | |||
| 2 | Two to three reference customers in production at our scale, called by us | Pass / Fail | |||
| 3 | Channel protection in writing: no direct sales or quotes to our clients | Pass / Fail | |||
| 4 | Integration tested in a sandbox before signature: one alert, one ticket, no duplicate loops, no second portal | Pass / Fail | |||
| 5 | Exit terms in writing: notice, export format, offboarding cost, data deletion | Pass / Fail | |||
| 6 | Liability and subprocessor terms in writing, including client-facing disclosure language | Pass / Fail | |||
| 7 | Contract term 12 months or less (36-month lock-in fails unless pricing is exceptional) | Pass / Fail |
Stage two: weighted score. Survivors only. Score 0 to 5.
| Criterion | Weight | Score (0 to 5) | Weighted | Note |
|---|---|---|---|---|
| Price at our volume | 3 | |||
| Cost fit (margin fit if you resell) | 3 | |||
| Billing fit | 2 | |||
| Price scalability at 2x volume | 2 | |||
| Peer and reference signal | 2 | |||
| Security posture and certifications | 2 | |||
| Support quality and responsiveness | 2 | |||
| Data location and sovereignty | 1 | |||
| Integration with the existing stack | 3 | |||
| Scope (what they take off our plate) | 3 | |||
| Automation depth | 3 | |||
| Prevention alignment | 2 | |||
| Management overhead added | 2 | |||
| Onboarding, migration and training | 2 | |||
| Total | 32 | / 160 |
Rebuilding this vendor scorecard template in Excel
Two formulas do all the work.
Paste both tables onto one sheet, gate first, score below it. Using the column layout above, the gate verdicts land in C2:C8 and the scoring block starts at row 12.
For the gate, put the result in C9:
=IF(COUNTIF(C2:C8,"Fail")>0,"STOP","PROCEED")
Conditional-format that cell red on STOP.
For the score, weights in B12:B25 and scores in C12:C25:
=SUMPRODUCT(B12:B25,C12:C25)
SUMPRODUCT multiplies each weight by its score and sums the results in one pass, so you never need a helper column. For the percentage, divide by the maximum rather than hardcoding 160, so the sheet survives you changing a weight:
=SUMPRODUCT(B12:B25,C12:C25)/(SUM(B12:B25)*5)
Then wire the two together, so the scoring block cannot produce a number while the gate says STOP:
=IF($C$9="STOP","GATE FAILED",SUMPRODUCT(B12:B25,C12:C25))
That last one is the whole point of building it in a spreadsheet rather than a document. The score is unreachable until the gate clears, so the sheet cannot do the thing this article is about.
Add data validation on C12:C25 restricted to whole numbers between 0 and 5, so nobody types 4.5 when they mean "I am not sure". Freeze the criterion column, one vendor per column pair, every vendor on one tab. Splitting vendors across tabs is how a comparison stops being a comparison.
Worked example: the AI SOC category
The hardest live application of this instrument right now is AI SOC tools, so let me run the gate against the category rather than any individual company. Everything below is public and dated.
For the vendor-by-vendor detail behind this category, see the AI SOC companies comparison and the five buyer tests in AI SOC software vs MDR.
It is the right test case because the claims are enormous, the money behind them is real, and the category is new enough that almost none of the usual proof exists yet. Every vendor in it will tell you its agents close most of your tier-1 work. Ask who checked, and the room goes quiet. That is exactly the gap a gate is built to find.
Now the vendors. Everything below came off their own public pages and funding coverage on 13 August 2026, and where I say something is absent I mean I went looking and did not find it published, which is not the same as it not existing. That distinction matters more than usual here, because a gate turns on evidence you can hold rather than on what a vendor may have privately.
| Vendor | Founded | Raised | Staff | Published certifications |
|---|---|---|---|---|
| Exaforce | 2023 | about $200M | about 110 | SOC 2 Type II, ISO 27001, HITRUST |
| AirMDR | 2023 | $15.5M | about 38 | None published on its site as of 13 August 2026 |
| 7AI | 2024 | about $166M | about 125 | SOC 2 Type II |
| TENEX.AI | 2024 | about $277M | about 132 | SOC 2 Type II, HIPAA, PCI-DSS |
Torq belongs in the conversation but not in that table, since it is an automation platform that has moved into the category rather than an AI-native startup inside it. It publishes SOC 2 Type II and ISO 27001. A few more public specifics on all five, because the detail is where the gate bites.
Torq calls itself an AI SOC platform, and its product Socrates an AI SOC analyst. Its managed-services page says Socrates "autonomously handles Tier-1 and Tier-2 cases and can act as a 24x7 on-call agent". It is agentless and API-first, CrowdStrike is a named integration, and its named reference partners are themselves MDR and MSSP providers running Torq internally: Deepwatch, RSM Defense, HWG Sababa. Pricing is quote-gated (the /pricing page returns a 404) and modelled per workflow and automation volume rather than per user or endpoint.
On Exaforce's own integrations page, the CrowdStrike, SentinelOne, Office365, ServiceNow and Jira entries are all marked "Coming soon". TechCrunch reported roughly 20 customers as of May 2026.
AirMDR's founder co-founded Sumo Logic. It sells two paid tiers: "Full Service AI MDR" marketed "For Small Security Teams" and staffed by AirMDR, and "AI SOC Platform" marketed "For MSSP and Enterprise SOC Teams that Need AI Analysts", which the customer runs. Its CrowdStrike integration is live as a data source and for response actions through the customer's own Falcon deployment. Zero reviews on G2, zero on PeerSpot.
7AI was founded by the former CEO and CTO of Cybereason and positions as "service as software". On the default product the customer's own analysts stay on the loop, and its staffed tier, PLAID ELITE, only launched in May 2026. Its flagship channel deal is with DXC, a $14bn systems integrator that embeds 7AI inside its own SOC. No published pricing, zero reviews on G2.
TENEX.AI raised at a valuation above $1bn. Its flagship tier is staffed: "24/7/365 case management by TENEX", "100% named analysts, 8+ yrs average", "zero offshore handoffs". Its homepage says "Agents own 95%+ of T1 and T2". It has a dedicated CrowdStrike page covering management of a customer's existing Falcon deployment including containment, and prices "on ingestion, not per token or per head". One named customer, Sunrun. No MSP partner program located.
| Vendor | Founded | Raised | Staff | Independent proof of AI claims |
|---|---|---|---|---|
| Exaforce | 2023 | $200m | ~110 | none found |
| AirMDR | 2023 | $15.5m | ~38 | none found |
| 7AI | 2024 | $166m | ~125 | none found |
| TENEX.AI | 2024 | $277m | ~132 | none found |
Every company here was founded in 2023 or 2024, and every performance number I could find is reported by the company selling it. Gate item 1 stops the field on its own.
Here is the pattern. Every one of the AI-native names was founded in 2023 or 2024. Across all of them I could not find independent third-party validation of the AI or automation claims, which means every performance number I could find is reported by the company selling it. I could not find a published named reference customer at mid-market scale for any of them either.
Which means that on public evidence, gate items 1 and 2 stop every one of them today. Not one company. All four.
That is the instrument working correctly, and it is not a criticism of these businesses. They are young, well funded and building fast, and several will be obvious buys in two years. The gate is telling you something narrower: the evidence a careful buyer needs is not public yet, so unless a vendor produces it privately, you are buying on faith.
There is a buyer's version of the hype cycle worth holding onto. At the peak of attention you are the experimental budget, and the trough is where weak vendors get shaken out. Survivable when the thing you piloted was a marketing tool. When the vendor holds standing access to your customers' systems, a vendor failing is an incident, not a wasted pilot. That asymmetry is why the gate exists, and it is the same reasoning behind an AI risk assessment template and a written AI acceptable use policy: decide the unacceptable cases while nobody is selling to you.
The layer test: who employs the humans
One question cuts through most of the confusion in this category, and it is not technical.
Who employs the humans who answer at 2am?
Torq's "24x7 on-call agent" is software. The marketing copy says so, and it is a legitimate product, but it is a licence you operate, and if it misjudges a case at 2am the person who picks it up works for you. AirMDR sells both shapes at once: its channel tier hands an MSSP the platform to staff itself, while its staffed tier points at small in-house security teams. Same brand, two different things being bought.
In AI-era categories the language can blur this, because "AI analyst" sounds like a person and prices like software. So make the layer explicit before you score anything. If the vendor's own employees carry the pager, you are buying a service, and support quality, escalation paths and analyst tenure are the criteria that matter. If your team carries it, you are buying a licence, and management overhead added is suddenly your most important criterion rather than a weight-2 afterthought.
The related trap is alert forwarding described as monitoring. A vendor that watches a console and emails you when something lights up has moved the work by one inch and charged you for a SOC. Ask what happens after the alert. If the answer is that you get notified, you have bought a notification. This is the same layer confusion that decides who keeps the efficiency gain in AI SOC economics, and you want to know which side of it you are standing on before the pricing conversation starts.
What the scorecard does not fix
Back to the team who never used the thing.
The instrument was good. It is the one on this page. What it could not do was survive a room where people already had opinions, a vendor they liked, and a meeting to get through. Frameworks lose to rooms, reliably, and being correct makes no difference. A scorecard is a discipline with a spreadsheet attached. Ship the spreadsheet without the discipline and you have shipped nothing.
Two things make it stick, and they are both boring.
Every gate item has a named owner and a date. Not the team, not procurement, a person, with a date the answer is due. That is why the template has Owner and Date columns. An unowned gate item is a wish. An owned one produces either an answer or a visible failure to get one, and both are useful.
The gate is answered before the demo, not after. This is the whole game. Once a team has watched a good demo, the gate stops being a filter and becomes an obstacle between them and a decision they have already made, and they will find a way around it. Answered first, the same seven questions cost an hour and often end the process before anyone has spent a day on it. The order matters more than the content.
The rest is honesty about your own zeros. A scorecard full of confident middling numbers is one where nobody wanted to write down how little they knew.
None of this makes the decision for you. It makes the decision visible, which is a lower bar and a more useful one. Same principle as writing down how AI gets governed inside your own business before a client asks, or being clear-eyed about what actually defends your position as a service provider. The value sits in having decided in advance, in writing, while it was still cheap.
FAQ
Two stages. A pass/fail gate of non-negotiables where any single Fail ends the evaluation, and a weighted 0 to 5 scoring block that only survivors reach. Most templates ship only the second half, which is why a vendor with a disqualifying flaw can still post a respectable weighted average and stay in the running.
Weights in one column, scores in the next, then =SUMPRODUCT(B12:B25,C12:C25) for the weighted total and =SUMPRODUCT(B12:B25,C12:C25)/(SUM(B12:B25)*5) for the percentage. Add =IF(COUNTIF(C2:C8,"Fail")>0,"STOP","PROCEED") for the gate, and wrap the score formula so it returns "GATE FAILED" while the gate says STOP.
Independent validation of whatever capability you are actually buying, and reference customers at your scale that you call yourself. Those two are the gate items that eliminate most of a young category. After that, exit terms in writing, because they are the cheapest thing to negotiate before signature and the most expensive thing to negotiate after.
Gate as many as you like, since the gate is fast and mostly answerable from public material and one commercial call. Score no more than three or four. Scoring is slow, it needs the same people in one sitting, and a fifth vendor rarely changes the answer. If more than four survive the gate, your gate is too soft.
Then either the gate item was never truly non-negotiable, and you should rewrite it before your next evaluation, or you are about to make an exception you will meet again in month four. Both happen. What you cannot do is quietly convert the Fail into a low score and let the average absorb it, which is exactly the failure mode the gate exists to prevent.
Both. The instrument is the same and only two lines change wording. If you resell, gate item 3 is channel protection and the cost criterion is margin fit. If you buy for your own estate, gate item 3 is the vendor agreeing not to sell around IT into your business units, and the cost criterion is whether it fits the budget line all in, including the internal hours to run it. Everything else, including the attestation and exit questions, reads identically.
It works particularly well there, because the category is young enough that the gate does most of the sorting. Applied to the AI-native vendors above in August 2026, gate items 1 and 2 stopped all of them for me, because I could not find third-party validation of the AI claims or a published named mid-market reference. That is a finding about what is public today, not a verdict on any company.