Projects

Selected research and software projects spanning antibody engineering, interpretable AI, and scientific computing. Each entry states the problem, the approach, the measured result, and my role.

ProteinTuneRL

ProteinTuneRL logo

Problem. Therapeutic antibodies must be optimized for several measured biophysical properties at once — binding affinity, immunogenicity, expression — and a sequence’s likelihood under a pretrained protein language model is a poor proxy for any of them.
Approach. Reinforcement-learning post-training of an infilling language model (IgLM) that redesigns CDR loops in the context of the surrounding framework. Online path: PPO with KL regularization against the reference model. Offline path: Direct Reward Optimization adapted to proteins by adding a value head, so the model learns from static (prompt, response, feedback) assay data under an MSE objective — the common case for historical experimental campaigns.
Result. Improves alignment with measured biophysical properties and outperforms likelihood-only baselines across binding-affinity, immunogenicity, and expression tasks. KL regularization proved necessary: without it, score optimization destroys sequence plausibility.
My role. Project lead; senior author on the paper.
Stack. PyTorch · Hugging Face Transformers · IgLM · PPO · REINFORCE · DRO · Python 3.9+
Artifacts. GitHub · Paper: Reinforcement Learning for Antibody Sequence Infilling · November 2024 – Present


ProtLib-Designer

Protlib Designer logo

Problem. Designing an antibody library is a constrained combinatorial search: the batch must be high-affinity, developable, and diverse. Greedy selection collapses onto near-identical high scorers, and evolutionary search cannot guarantee it returns the library size you asked for.
Approach. In-silico deep mutational scanning scores every single-point mutation with a protein language model (ProtBert) and an antibody-specific inverse-folding model (AntiFold); those scores seed a multi-objective integer linear program in which diversity limits and per-sequence mutation budgets are hard constraints. An iterative solve-and-remove strategy scales the formulation to large libraries.
Result. On cold-start design for Trastuzumab, D44.1, and Spesolimab: ~3.5× the fraction of predicted binders of the greedy (LMG) and MODIFY baselines while also improving average humanness; highest hypervolume and best average rank across all metrics; returns the full 1,000 unique sequences by construction, where the SPEA2 evolutionary baseline came up ~26% short on D44.1.
My role. Project lead; senior author on the paper.
Stack. Python 3.10 · integer linear programming · ProtBert · AntiFold · published to PyPI with CI and a documented CLI
Artifacts. PyPI · GitHub · Paper: AAMAS 2026 · Project post · November 2024 – Present


Deep Symbolic Optimization

DSO banner

Problem. Deep models are accurate but opaque. Scientific discovery and safety-critical control need a closed-form expression or policy that a domain expert can read, check, and sign off on.
Approach. Treat discovery as sequential decision-making: an autoregressive generator emits expression trees and is trained with risk-seeking policy gradients, which optimize the best-case rather than the average sample. The unified framework (uDSR) combines neural-guided search with genetic programming and complementary strategies, and a task interface extends the same machinery from symbolic regression to symbolic control.
Result. State of the art on the SRBench benchmark in both symbolic-solution and accuracy-solution rate, and 1st place in the real-world track of the 2022 SRBench competition at GECCO. On control, decision-tree and closed-form policies that rival neural policies while remaining fully auditable. The codebase supports 11 publications, including an ICLR oral, an ICML spotlight, and two NeurIPS papers.
My role. First author of the unified framework (NeurIPS 2022) and of the symbolic-control method (ICML 2021 spotlight); second author on the original risk-seeking DSR paper (ICLR 2021 oral); senior author on DisCo-DSO (AAAI 2025); core contributor to the open-source codebase.
Stack. TensorFlow (PyTorch refactor available) · multiprocessing-parallel batch runs · pip-installable package with a task plugin API
Artifacts. GitHub · NeurIPS 2022 · ICML 2021 · AAAI 2025 · January 2021 – Present


Cardiac Machine Learning

Cardiac ML banner

Problem. Electroanatomical mapping of the heart normally requires an invasive catheter procedure. The standard 12-lead ECG is non-invasive and cheap, but low-dimensional — 12 channels standing in for a whole organ.
Approach. Generated 16,140 organ-level cardiac simulations on LLNL’s Lassen supercomputer (200 µm meshes, 5 µs time steps, 4 GPUs and 40 CPU cores per concurrent run), then trained 1D SqueezeNet-style CNNs (~0.4–0.5M parameters) to map a 12×500 ECG tensor to 75-site activation maps and to full transmembrane-voltage traces.
Result. Mean activation-time error of 1.66 ms across 75 intracardiac sites, and Pearson r = 0.97 for transmembrane-voltage reconstruction — accurate enough to recover septal and transmural activation patterns non-invasively. Dataset and training code released publicly; the approach is covered by a US patent.
My role. First author; co-curator of the released dataset; author of the open-sourced code and of the teaching challenge built on it; named inventor on the patent.
Stack. PyTorch · NumPy · h5py · Cardioid (C++/MPI/CUDA) for data generation · LLNL Lassen GPUs
Artifacts. GitHub · Cardiac Challenge · CinC 2022 paper · Dataset · US patent · Write-up · July 2018 – July 2021


Cardioid

Problem. Simulating a beating heart end to end — subcellular ion channels through tissue electrophysiology and mechanics to the surface ECG — at resolutions fine enough to be clinically meaningful, which puts the problem firmly on supercomputers.
Approach. Finite-element electrophysiology and cardiac mechanics on realistic anatomical meshes, with the electromechanical coupling in the left ventricle extended to include the Purkinje conduction network, plus the ECG forward problem and meshing and fiber-generation tooling.
Result. A production multiscale simulation suite, open-sourced by LLNL, that runs distributed and GPU-accelerated. It is also the engine that generated the 16,140-simulation dataset behind the cardiac ML work above.
My role. Contributor to the suite; first author on the electromechanical-coupling and Purkinje-network study.
Stack. C99/C++ · MPI · OpenMP · CUDA (NVTX/NVRTC) · MFEM · LAPACK · Spack builds
Artifacts. GitHub · Paper: Numerical approximation of the electromechanical coupling in the left ventricle with inclusion of the Purkinje network · 2018 – 2020


EXIFSI

Problem. In incompressible fluid-structure interaction, explicit (non-iterative) coupling is cheap but goes unstable through the added-mass effect; implicit coupling is stable but pays for it with an iteration loop at every time step. Biomechanics needs both stability and speed.
Approach. Robin-Neumann splitting schemes that make explicit coupling stable, and Nitsche-XFEM unfitted-mesh methods for fluid coupled to immersed thin-walled structures, each with stability and convergence analysis rather than empirical tuning.
Result. Non-iterative coupling schemes with proven stability, published in CMAME and JCP, and validated against the FSI forward-prediction challenge benchmark. The resulting thesis won the SMAI-GAMNI award for the best French thesis in numerical methods for the mechanical and engineering sciences, and was an ECCOMAS PhD award finalist.
My role. Ph.D. researcher on the 4-year ANR project; sole author of the thesis; first author on the forward-prediction validation study; co-author on the Nitsche-XFEM method.
Stack. C++ · FELiScE finite-element library · MPI · PETSc
Artifacts. Project website · CMAME 2016 · JCP 2015 · August 2012 – July 2016 · Funded by ANR JCJC (French National Research Agency)