How it works

How the shortlist is made

A shortlist with evidence, not a regulatory determination. Everything below is deterministic: no language model ranks, filters or writes any of it.

Data

Every 510(k) summary FDA has posted — 92,098 cleared devices across 3,098 product codes, 80,356 of them with a summary text, as of the last data run, read by fda510k-core: registry fields, the 807.92 sections, comparison tables and the citations each summary makes, labelled predicate / reference / other with a rule-based classifier measured at 1.00 accuracy on a 73-citation held-out set. Recalls come from openFDA's device-recall dataset. CBER's BK-numbered 510(k)s are out of scope.

Last data run: seed, started 2026-09-25 19:58 UTC, finished 2026-09-25 20:25 UTC, 0 errors.

Vocabularies

Indication tags are phrases mined per product code from the indications texts: a phrase becomes a tag when at least 3% of the code's summaries (and at least 3) use it and it stands as a complete list item somewhere; fragments fold into fuller phrases; words that recur across most product codes are treated as boilerplate. Codes with fewer than 20 summaries borrow the vocabulary of the neighbour code that cites them most. Attributes are comparison-table row labels plus materials, additives and constructions named in the description or the table's material row, with the role each plays (base, additive, coating).

Predicate score

WeightSignalWhat it measures
40%indication coveragethe share of your ticked indications the device carries
30%attribute matchthe share of your named attributes it carries (with the value, when you gave one)
15%description similaritycosine similarity of gte-small text fingerprints; neutral when unavailable
10%recency1.0 within 10 years, 0.75 within 15, 0.45 older
5%testing paritythe share of the testing kinds you will provide that its summary reports

Penalties: −0.1 for a quality-related recall on file, −0.15 when the device is outside your chosen codes. A design-related recall excludes a device from the predicate list (FDA's fourth best practice); it is still shown, marked excluded. Candidates are the devices with summaries in your codes plus neighbour codes linked by at least 3 citations, capped at 25,000.

Reference devices

For each attribute you named that the predicate lacks, devices carrying it are scored 35% for the attribute, 20% when a device in your codes already cites them, 20% indication coverage, 15% recency and 10% when the attribute is a minor share (coating, layer, additive) and your subject's share is 25% or less; −0.05 for a quality recall.

Known limits

Source: predicate-advisor (this site and its build) and fda510k-core (the corpus). FDA's predicate-selection best practices are cited from the draft guidance of September 2023.