Now Hiring: Founding Lead Researcher – Medical AI Evaluations | Fully Remote | $100K–$300K USD
Overview
Sophont is hiring a Founding Lead Researcher for Medical AI Evaluations, a fully remote leadership post focused on building the evaluation systems and standards used to assess medical AI models. Sophont is a public-benefit corporation developing open, universal medical AI that covers areas such as pathology, neuroimaging and clinical text, and works with MedARC, its open-source medical AI research community. Its open work has appeared in venues including NeurIPS, ICML, CVPR and Nature Biomedical Engineering, and it describes its environment as closer to a live research laboratory than a conventional startup.
The researcher owns the evaluation layer across the portfolio, including language models, pathology and neuroimaging models, multimodal and agentic systems, retrieval tools and clinical-facing products. Work builds on projects such as Medmarks, Brainmarks, Nanopath and OpenMidnight. Duties include building benchmark and leaderboard infrastructure, reproducible evaluation jobs, evaluation datasets with clinicians and community contributors, annotation rubrics, quality control and leakage checks, and standards for source-grounded systems covering retrieval quality, evidence recall and precision, uncertainty calibration and medical fact verification. The researcher also runs statistically rigorous comparisons, subgroup and robustness analyses, reports negative results clearly, writes papers and reports, and leads MedARC contributors.
Who can apply
Sophont wants an exceptional researcher experienced in machine learning evaluation, benchmarking, model validation or measurement infrastructure, with a strong grounding in statistics, experimental design and data quality, and strong Python and software engineering skills. Experience with medical, clinical, pathology or neuroimaging datasets, retrieval-augmented generation, LLM-as-a-judge systems, public benchmarks, crowdsourcing platforms or GPU and distributed evaluation is valuable, as is a publication or open-source record.
What you get
The annual salary ranges from $100,000 to $300,000 USD, with meaningful equity, OpenAI and Anthropic subscriptions, a 401(k) with a 4% match, and medical, dental, vision and basic life insurance. Staff largely set their own hours around core collaboration meetings and are encouraged to publish and take part in conferences.
How to apply
Candidates apply by sending their CV by email using a subject line made of their full name, then CV, then Medical AI Evaluations.