How it works
How the shortlist is made
A shortlist with evidence, not a regulatory determination. Everything below is deterministic: no language model ranks, filters or writes any of it.
Data
Every 510(k) summary FDA has posted — 92,098 cleared devices across 3,098 product codes, 80,356 of them with a summary text, as of the last data run, read by fda510k-core: registry fields, the 807.92 sections, comparison tables and the citations each summary makes, labelled predicate / reference / other with a rule-based classifier measured at 1.00 accuracy on a 73-citation held-out set. Recalls come from openFDA's device-recall dataset. CBER's BK-numbered 510(k)s are out of scope.
Last data run: seed, started 2026-09-25 19:58 UTC, finished 2026-09-25 20:25 UTC, 0 errors.
Vocabularies
Indication tags are phrases mined per product code from the indications texts: a phrase becomes a tag when at least 3% of the code's summaries (and at least 3) use it and it stands as a complete list item somewhere; fragments fold into fuller phrases; words that recur across most product codes are treated as boilerplate. Codes with fewer than 20 summaries borrow the vocabulary of the neighbour code that cites them most. Attributes are comparison-table row labels plus materials, additives and constructions named in the description or the table's material row, with the role each plays (base, additive, coating).
Predicate score
| Weight | Signal | What it measures |
|---|---|---|
| 40% | indication coverage | the share of your ticked indications the device carries |
| 30% | attribute match | the share of your named attributes it carries (with the value, when you gave one) |
| 15% | description similarity | cosine similarity of gte-small text fingerprints; neutral when unavailable |
| 10% | recency | 1.0 within 10 years, 0.75 within 15, 0.45 older |
| 5% | testing parity | the share of the testing kinds you will provide that its summary reports |
Penalties: −0.1 for a quality-related recall on file, −0.15 when the device is outside your chosen codes. A design-related recall excludes a device from the predicate list (FDA's fourth best practice); it is still shown, marked excluded. Candidates are the devices with summaries in your codes plus neighbour codes linked by at least 3 citations, capped at 25,000.
Reference devices
For each attribute you named that the predicate lacks, devices carrying it are scored 35% for the attribute, 20% when a device in your codes already cites them, 20% indication coverage, 15% recency and 10% when the attribute is a minor share (coating, layer, additive) and your subject's share is 25% or less; −0.05 for a quality recall.
Known limits
- Attributes and tags are read automatically; older scanned summaries yield fewer. Cards say when a device's text was thin.
- Comparison tables exist mostly from 2011 on.
- openFDA updates weekly and FDA posts summary PDFs days to weeks after a decision, so the newest clearances appear first as registry-only cards and gain their summary when the daily job finds it posted.
- Predicate/reference labels are rule-based with measured 1.00 accuracy on a 73-citation held-out set; unclassified citations are shown as such, not guessed.
Source: predicate-advisor (this site and its build) and fda510k-core (the corpus). FDA's predicate-selection best practices are cited from the draft guidance of September 2023.