Skip to content

The method

How matching works

TrialMatch is a hybrid system. Objective comparisons are made by deterministic rules, not by a language model. AI is used for retrieval, for reading messy criteria text, and for wording the explanation — never for deciding whether you are eligible.

Six stages, in order

  1. Stage 1

    Data synchronisation

    We consume the official ClinicalTrials.gov API, never the public website. Records are stored locally with the registry's own last-updated timestamp and a content hash, so we can tell a genuine change from a re-publish. Your searches always run against this local mirror, so a registry outage never blocks you.

  2. Stage 2

    Candidate retrieval

    Your structured profile drives a filtered search: condition, age, sex, recruitment status, study type and travel radius. Semantic similarity then re-ranks what survives, so that "stage IV lung cancer" can still find a study written up as "advanced metastatic NSCLC".

  3. Stage 3

    Hard filtering

    Studies you could not join are removed before scoring: wrong recruitment status, an age or sex limit you fall outside, an unrelated condition, or nothing running within the distance you set.

  4. Stage 4

    Criterion-level evaluation

    Free-text eligibility criteria are parsed into structured rules that keep a pointer back to the exact sentence they came from. A rule engine written in plain PHP compares each one against your information and returns Match, Mismatch or Unknown. Nothing here involves a model.

  5. Stage 5

    Score and status

    Five factors combine into a 0-100 match score: condition relevance (30), structured criteria met (30), semantic similarity (20), location fit (10) and profile completeness (10). A known major exclusion overrides the number entirely and drops the result to Likely Not Eligible.

  6. Stage 6

    Explanation

    The explanation is generated from the evaluation output alone. It cannot introduce a criterion, a location or a contact that is not already in the record, and any generated text that claims confirmed eligibility is rejected before you ever see it.

Published, so you can check it

The scoring weights

Each factor has a ceiling. Together they add up to 100, and you can see exactly what moves a result.

Condition relevance
30 pts

How closely the study subject matches your recorded diagnosis.

Structured criteria met
30 pts

The share of machine-readable criteria your information satisfies, reduced by how many stayed unknown.

Semantic similarity
20 pts

Conceptual closeness between your profile text and the study text. A ranking signal only.

Location fit
10 pts

Distance to the nearest recruiting site, against the travel limit you set.

Profile completeness
10 pts

How much of your profile is filled in. More information means a more confident result.

Honest answers

Why unknown is its own answer

A screening tool that forces every criterion into yes or no will quietly turn "I do not know" into "no". That is how people get told they are ineligible for a study they could have joined, or the reverse. We keep unknown as a first-class state everywhere: in the database, in the score, in the explanation, and on screen. Unknown criteria lower confidence in a result — they are never counted as either a match or a mismatch.

What this tool will not do
  • Match

    Your information appears to satisfy the criterion.

  • Mismatch

    Your information appears to conflict with the criterion.

  • Unknown

    There is not enough information to evaluate it. Never silently treated as a match or a mismatch.

Please keep in mind: This platform provides information about clinical trials and estimates potential eligibility based on the information you provide. A match does not mean that you qualify for a clinical trial. Final eligibility is determined by the study team. This service does not provide medical diagnosis or treatment advice.