AfterQuery · Research Scientist - Post Training

Prove whether the data caused a real capability shift.

Mohamed A M Elansary, PhD — multimodel evaluation under uncertainty, dataset-impact measurement, and production agent evaluation sets for SFT and RL post-training experiments.

Post-training evaluationDataset-impact measurementUncertainty quantificationProduction agent evals

Evaluation under uncertainty

  • Six-plus years of multimodel, multi-basin forecast experiments across hydroclimates on Linux/HPC.
  • Compared statistical and physically based stacks, quantified uncertainty, and reported regime-dependent failure modes rather than a single flattering score.
  • That is the measurement analogue of “this data → this improvement → under these conditions.”

Agents and public-eval transfer

  • Production GPT, Claude, and Gemini agent workflows with retrieval, routing, tenant isolation, provenance, and regression evaluation sets at Vertexium.
  • That maps to building evaluation sets and inspecting tool-using trajectories. It is not AfterQuery SFT/RL ownership.
  • Comfort reporting messy results clearly — the posting’s bias toward building over theorizing, applied to measurement.

Proposed first contribution

For one AfterQuery dataset family already used in an SFT or RL post-training experiment, define the intended capability claim and a small failure taxonomy: metric movement without a capability change, slice-specific collapse, generalization loss, alignment regression under a named condition. Stand up a small evaluation set with provenance on traces and data lineage, compare simple baselines, attach uncertainty, and write a clear external-facing report before expanding the training loop. This is a proposed measurement approach, not a claim of prior AfterQuery-internal work, SFT/RL ownership, or invented metrics.

Honest fit boundary

Two gaps are stated plainly. The live preferred-qualifications line says great candidates are undergrad or master’s research and have not done a PhD; I have a PhD and do not hide that preference mismatch. Direct post-training and RL depth is also a stretch: I have not run SFT or RL post-training loops, trained reward models, or claimed AfterQuery-internal work, and I do not invent metrics or safety research. The credible contribution is evaluation under uncertainty, measurement of messy experiments, scientific/HPC rigor, and production agent evaluation harnesses.

Role and location

Research Scientist - Post Training · San Francisco · OnSite. Ashby lists workplaceType OnSite and isRemote false. Willing to relocate to San Francisco with a relocation package. Remote eligibility is not asserted.

Posting compensation: “$210K – $450K • Offers Equity • Offers Bonus • $250k-450k total compensation + equity”. · Official role posting