Skip to content
View tomaszgy's full-sized avatar

Block or report tomaszgy

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
tomaszgy/README.md

Hi, I'm Tomasz Gargas 👋

Staff-level ML Engineer with 12 years of experience across Knowledge Graphs, Proteomics, and Generative Biology.

I build end-to-end discovery platforms for high-dimensional biological data. My focus is on architecting the end-to-end framework —from initial hypothesis modeling to production-grade discovery engines—ensuring that high-scale biological insights remain both reproducible and biologically consistent.

🛠 Technical Expertise

  • GenBio & scRNA-seq: Scaling and benchmarking Foundation Models (scGPT, Geneformer) for cell-type annotation and gene perturbation.
  • Agentic Systems: Designing stateful multi-agent RAG and dialectic reasoning loops via LangGraph (see Cherchoux).
  • Proteomics & MTL: Architecting Multi-Task Learning frameworks for population-scale proteomics (UKB-PPP), optimizing for rare-phenotype signal retention.
  • Relational ML & Graphs: Large-scale Knowledge Graph modeling (Neo4j/CERN) and Link Prediction for therapeutic discovery.
  • Cloud & MLOps: High-autonomy contributor in GCP/AWS environments with a strict focus on reproducibility and data integrity from ingestion to inference.

🧬 Professional Impact (Proprietary)

  • scFM Benchmarking Library: Architected a unified evaluation framework to standardize the fine-tuning and embedding generation of single-cell FMs, reducing R&D experiment cycles from weeks to days.
  • Validated Link Prediction: Developed a relational ML pipeline for drug-target discovery. Results were selected for wet-lab validation, identifying novel therapeutic candidates.
  • UKB-PPP Multi-Task Engine: Scaled a proteomics embedding pipeline to handle high-dimensional, multi-phenotype population data.

🌐 Open Source & Leadership

  • Contributor: edx-platform (Python/Django) – contributed to infrastructure serving millions of learners globally.
  • CERN Alumnus: Designed large-scale bibliometric and citation graphs using Neo4j.
  • Remote-First Engineering: Over a decade of experience in distributed teams. I specialize in co-designing automated discovery workflows that enable rapid scientific iteration while maintaining the rigorous data integrity required for therapeutic discovery.

Pinned Loading

  1. cherchoux cherchoux Public

    Stateful multi-agent system for adversarial hypothesis testing in biomedical research.

    Python 1

  2. inspire-relations inspire-relations Public

    Forked from inspirehep/inspire-relations

    Invenio module to integrate Neo4J graph database into INSPIRE and handle relations across records.

    Python

  3. inspirehep/inspire-next inspirehep/inspire-next Public

    The INSPIRE repo.

    Python 60 71

  4. edx-platform edx-platform Public

    Forked from openedx/openedx-platform

    The Open edX platform, the software that powers edX!

    Python

  5. openedx/openedx-platform openedx/openedx-platform Public

    The Open edX LMS & Studio, powering education sites around the world!

    Python 8.2k 4.3k

  6. lemot lemot Public

    Personalized vocabulary acquisition engine using agentic context-generation and adaptive retention logic.

    Python