What is Hit-to-Lead and when do you need it?
Hit-to-Lead is the filtering work that sits between a screen producing actives and a chemistry team committing to optimize. A high-throughput or fragment screen hands you a list of compounds that lit up the assay. Most of that list is noise: aggregators, frequent hitters, assay interference (PAINS), reactive groups, and singletons with no follow-up chemistry. Hit-to-lead is where you confirm which actives are real, kill the artifacts, and narrow a long primary-hit list down to two or three chemical series clean enough to justify the expense of full lead optimization.
You need it the moment you have hits but no defensible reason to pick one over another. The concrete work is hit confirmation by re-test and orthogonal assay, full dose-response to get real IC50 or EC50 values rather than single-point activity, counter-screens and selectivity panels to weed out false positives, a structure check (LC-MS, NMR) to confirm the compound is what the plate label says, and the first rounds of structure-activity relationship (SAR) to see whether a series actually responds to chemistry. Early developability flags get pulled in here too: kinetic solubility, a microsomal stability read, sometimes an early hERG or cytotoxicity screen, so you do not optimize a series with a fatal liability baked in.
The reason hit-to-lead exists as its own decision gate is economic. Lead optimization is the expensive, multi-month design-make-test engine, and you only want to point it at series that can win. A disciplined hit-to-lead phase is what stops a program from spending a year optimizing a frequent hitter or a chemically dead series. The output is not a drug. It is a ranked, de-risked set of validated hit series with enough early SAR and developability data to choose what to advance, and what to drop.
What does a Hit-to-Lead CRO actually do?
A hit-to-lead CRO takes your screening output and runs the triage-and-validation work, usually as a tightly scoped package with a defined endpoint rather than an open-ended chemistry engagement. The deliverable buyers care about is a confirmed, prioritized series list with the data behind each call.
Most providers do some combination of the work below. The mix depends on whether you are bringing a small-molecule screen, a fragment campaign, or a biologics hit panel, so match the scope to your modality before comparing quotes.
- Hit confirmation and validation: re-test of primary actives, orthogonal and label-free assays (SPR, ITC, thermal shift) to confirm real target binding and rule out assay-format artifacts.
- Dose-response and potency: full concentration-response curves for IC50 or EC50, replacing single-point screen data with numbers you can rank a series on.
- False-positive triage: removing aggregators, PAINS, redox-active and reactive compounds, frequent hitters, and singletons with no synthetic follow-up.
- Counter-screens and selectivity: testing against related targets, off-target panels, and a cytotoxicity control so apparent potency is real, not a general cell-health effect.
- Identity and purity: LC-MS and NMR confirmation that the active compound is the intended structure at adequate purity, plus re-supply or re-synthesis of confirmed hits.
- Early SAR and series triage: limited analog synthesis around each cluster to test whether activity tracks structure, then clustering hits into discrete chemical series and ranking them.
- Early developability flags: kinetic solubility, microsomal or hepatocyte stability, ligand efficiency and lipophilic efficiency, and sometimes an early hERG or genotoxicity flag to retire liabilities before optimization.
How to choose a Hit-to-Lead CRO?
The first filter is fit to your chemistry and modality, not the size of the logo. A CRO that excels at small-molecule HTS triage is not automatically the right shop for a fragment campaign needing crystallography and biophysics, or for a biologics hit panel. Ask for case studies in your target class, and check that the medicinal chemist who will rank your series has done this kind of triage before. Past that, work the checklist below, and score two or three candidates against the same written scope so the quotes measure the same thing.
- Quality and GxP status: hit-to-lead is research-grade work under good scientific practice, not GLP, so look for documented assay qualification (Z-prime, reproducibility), ISO 9001 or an equivalent quality system, and clean electronic-notebook practice rather than a GLP certificate you do not need here.
- Capacity and lead time: confirm the assay and biophysics platforms are running now, not being stood up on your dollar, and pin down turnaround for the confirmation round and each early-SAR cycle, since a slow cycle stacks across the iterations a real triage needs.
- Modality and indication fit: small molecule, fragment, peptide, or biologic each carry a different assay menu and a different definition of a clean series, so match the supplier to what you are actually screening.
- Region and regulatory track record: hit-to-lead rarely touches a regulator directly, but confirm where the work is run and that data capture is traceable enough to support the IND-enabling package that comes later.
- Data quality and reporting: insist on clean SAR tables, full dose-response curves, honest reporting of the compounds and series that failed, and defined assay acceptance criteria, because a cheap triage you cannot trust sends you into optimization blind.
- IP and confidentiality: settle who owns confirmed hits, analogs, and any platform-derived inventions before work starts, and confirm how compounds and data transfer to you, plus how the supplier protects a target you may not want disclosed.