How tots works

A CVE is a statement, not a verdict. Scanners repeat it, maintainers argue with it, and the version range drifts between databases. Given one CVE and one package version, tots asks the two questions a security engineer would: does the code actually do this? and what do the people who own it say?

New to security, or want the background? See how a CVE goes from discovery to disclosure, dispute, and patch, plus a glossary of every acronym on this page.Anatomy of a CVE →

The pipeline

One durable eve workflow runs every assessment. Each box is a step that survives restarts: if a model call fails halfway, the run resumes where it stopped instead of starting over.

01
Record
Fetch the CVE from NVD, the GitHub Advisory Database, and OSV with plain HTTP calls, with no model involved. Each source keeps its own version range, because disagreement between sources is evidence.
↓
02
Claims
A model breaks the description into atomic claims (product, each source's range, auth, vector, impact), plus the one the label hinges on: <package@version> is affected.
↓
03a
Technical investigator
Runs in an isolated Vercel Sandbox. Installs the target, the fixed version, and a known-vulnerable “positive control”, reads the patch, and runs one minimal PoC against all of them. It is blocked from reading issue threads, so opinions can't leak in.
03b
Discourse investigator
Reads advisories, GitHub issues and PRs, and vendor notes. Records who said what, with their role (using GitHub's own maintainer marker), a verbatim quote, and a link. No sentiment scores.

in parallel · neither sees the other's work

↓
04
Judge
Jev, an evaluation model, answers typed questions over both reports: is the target affected? Was it reproduced? Is there a credible dispute, and how authoritative is it? Each answer is a probability, not a paragraph.
↓
05
Policy
Fixed thresholds turn those probabilities into a label. No model picks the label, and every report shows the rule that fired.

The labels

  • INSUFFICIENT EVIDENCEChecked first: not enough concrete evidence to decide.
  • DISPUTEDA credible dispute that the PoC doesn't settle either way.
  • LIKELY INVALIDThe target looks unaffected: the PoC works on a vulnerable version but not the target, or maintainers dispute it.
  • LIKELY OVERSTATEDThe bug is real on the target, but the claimed impact isn't what the evidence shows.
  • SUPPORTEDReproduced on the target version, and no claim is credibly disputed.
  • LIKELY SUPPORTEDProbably affected, but not reproduced, with no strong dispute.

The official CVE status is always shown next to the label. The CVE Program has the final word; tots says how well the public evidence supports the claim for this version.

A real one

CVE-2024-10491express@5.2.1DISPUTED
Code
Reproduces on 5.2.1
no fixed release found
People
Maintainers dispute
3 say affected · 5 say not

→ DISPUTED: credible_dispute_exists 0.95 ≥ 0.60 and no stronger rule applied

This is the case tots exists for. Scanners flag it, the maintainers say it doesn't apply, and the sandbox shows the behaviour is still there. Neither side is simply right, and the report shows both.

Under the hood

Everything in tots is built on Vercel. The site, the agent, the durable workflow, the sandbox, every model call, the feature flag, and the database all live in one Vercel project. The Vercel services authenticate with the project's OIDC token, so there are no AI provider keys to manage.

Browser
↓HTTPS
Vercel · one project, one deployment
Next.js on Vercel
Pages, report UI, server actions that queue runs
Vercel Flags
live-runs: checked before any run is queued
Neon Postgres · Vercel Marketplace
Runs, stages, labels, full reports
↓/eve/v1 · same origin
eve agent service
Vercel's agent framework, deployed as its own service
↓
Vercel Workflow · investigate_cve
Durable steps: record → claims → investigators → judge → policy
↓in parallel
Technical subagent → Vercel Sandbox
Isolated microVM, egress limited to npm + GitHub
Discourse subagent
Advisories, GitHub issues, vendor notes
↓every model call
Vercel AI Gateway
Claude (investigators) · OpenAI (claims) · Jev (judge)
Vercel OIDC authenticates the Gateway, Sandbox, and Flags. No AI provider API keys.
↕public data: NVD · GitHub Advisories · OSV · GitHub issues · npm
  • eve
    Vercel's framework for durable agents: the workflow tool, both subagents, and their sandboxes
  • Vercel Workflow
    Each run is a durable workflow; a crash or redeploy resumes it mid-step
  • Vercel Sandbox
    Isolated microVMs for PoCs, egress limited to npm and GitHub
  • Vercel AI Gateway
    One endpoint for Claude, OpenAI, and Jev, billed and observed in one place
  • Vercel Flags
    Live runs are off by default; only a flag-checked server action can queue one
  • Neon via Vercel Marketplace
    Postgres provisioned from the Vercel dashboard, env vars injected
  • Next.js on Vercel
    This site, deployed alongside the agent with withEve
  • Vercel OIDC
    Short-lived project tokens for the Gateway, Sandbox, and Flags; no provider API keys

Where this could go

A threat-intel feed
A read-only API that any scanner, SBOM tool, or TIP can query before it raises an alert:
GET api.tots.dev/CVE-2024-10491?purl=pkg:npm/express@5.2.1

{
  "label": "DISPUTED",
  "code": "reproduces on 5.2.1",
  "maintainers": "dispute",
  "report": "https://tots-security.vercel.app/CVE-2024-10491/express@5.2.1"
}
VEX out of the box
Emit CycloneDX/OpenVEX statements from the label (LIKELY INVALID → not_affected with the PoC as justification), so the assessment travels with the SBOM.
In the pull request
A GitHub app that comments on Dependabot and scanner PRs with the tots label, so reviewers stop arguing about the same CVE in every repo.
Watch, don't snapshot
eve schedules re-check an assessment when a fix ships, an advisory changes range, or a maintainer weighs in, and flag when a label flips.