Prove whether the data caused a real capability shift.
Mohamed A M Elansary, PhD — multimodel evaluation under uncertainty, dataset-impact measurement, and production agent evaluation sets for SFT and RL post-training experiments.
Evaluation under uncertainty
- Six-plus years of multimodel, multi-basin forecast experiments across hydroclimates on Linux/HPC.
- Compared statistical and physically based stacks, quantified uncertainty, and reported regime-dependent failure modes rather than a single flattering score.
- That is the measurement analogue of “this data → this improvement → under these conditions.”
Agents and public-eval transfer
- Production GPT, Claude, and Gemini agent workflows with retrieval, routing, tenant isolation, provenance, and regression evaluation sets at Vertexium.
- That maps to building evaluation sets and inspecting tool-using trajectories. It is not AfterQuery SFT/RL ownership.
- Comfort reporting messy results clearly — the posting’s bias toward building over theorizing, applied to measurement.
Proposed first contribution
For one AfterQuery dataset family already used in an SFT or RL post-training experiment, define the intended capability claim and a small failure taxonomy: metric movement without a capability change, slice-specific collapse, generalization loss, alignment regression under a named condition. Stand up a small evaluation set with provenance on traces and data lineage, compare simple baselines, attach uncertainty, and write a clear external-facing report before expanding the training loop. This is a proposed measurement approach, not a claim of prior AfterQuery-internal work, SFT/RL ownership, or invented metrics.
Honest fit boundary
Two gaps are stated plainly. The live preferred-qualifications line says great candidates are undergrad or master’s research and have not done a PhD; I have a PhD and do not hide that preference mismatch. Direct post-training and RL depth is also a stretch: I have not run SFT or RL post-training loops, trained reward models, or claimed AfterQuery-internal work, and I do not invent metrics or safety research. The credible contribution is evaluation under uncertainty, measurement of messy experiments, scientific/HPC rigor, and production agent evaluation harnesses.
Role and location
Research Scientist - Post Training · San Francisco · OnSite. Ashby lists workplaceType OnSite and isRemote false. Willing to relocate to San Francisco with a relocation package. Remote eligibility is not asserted.
Posting compensation: “$210K – $450K • Offers Equity • Offers Bonus • $250k-450k total compensation + equity”. · Official role posting