ML systems · Tensor methods · Foundations

Yang Qi

Research Engineer

I build ML systems and develop algorithms for efficient learning, spanning language-model training and post-training, tensor methods, and mathematically grounded machine learning.

Portrait of Yang Qi

Areas of focus

  • Efficient ML architectures
  • LLM training & post-training
  • Tensor methods
  • Mathematical ML foundations

01 / Selected work

Selected Projects

Selected work across language-model systems, efficient tensor architectures, and high-dimensional learning theory.

Independent ML systems

PetitGPT

A 124.6M Language Model Trained from Scratch

A released language model pretrained on approximately 13B token positions using one RTX 4090, with a custom tokenizer, instruction tuning, reproducible evaluation, and native PyTorch inference.

  • Built a 30-layer, 124.6M-parameter decoder with grouped-query attention, RMSNorm, SwiGLU, and tied embeddings; pretraining reached reference validation loss 2.4702.
  • Reached 57.74% ARC-Easy and 28.16% ARC-Challenge accuracy, ahead of two evaluated SmolLM 135M instruct baselines under the same protocol; results were lower on PIQA and HellaSwag.
  • Published the weights, tokenizer, training recipes, evaluation protocols, loss curves, and success/failure cases; studied adaptation and capability retention through SFT, DPO, response distillation, and LoRA.
Model size
124.6Mreleased parameters
Pretraining
13Btoken positions
Training hardware
RTX 4090
Huawei · Efficient tensor learning

Efficient Tensor Approximation

Alternating CNN

A neural architecture for compressing high-dimensional CSI tensor data through efficient approximation of low-rank tensor structure.

  • Developed a neural architecture for high-dimensional CSI tensor data compression.
  • Designed the approach to efficiently approximate low-rank tensor structure.
  • Achieved normalized reconstruction error below 0.01 and reduced training time by approximately 25% relative to a 3D-CNN baseline.
Normalized reconstruction error
< 0.01achieved
Training time
≈25% lowerrelative to 3D-CNN baseline
Reference
3D-CNNbaseline
Huawei · Learning theory

Correlated Spiked Tensor Models

High-dimensional recovery

Theory and algorithms for recovering multiple correlated spikes in high-dimensional tensor models.

  • Studied detection phase transitions and local-optimization accuracy using random matrix theory and high-dimensional optimization.
  • Derived statistical guarantees for correlated latent components.
  • Derived an asymptotically unbiased estimator of signal strength.

02 / Experience

Experience

Recent machine learning work, supported by a longer record in mathematical and applied research.

Nov 2023 – Dec 2024

Paris, France

Research Scientist

Huawei Paris Research Center

  • Developed an efficient neural architecture for low-rank CSI tensor approximation, achieving normalized reconstruction error below 0.01 and reducing training time by approximately 25% relative to a 3D-CNN baseline.
  • Developed theory and algorithms for correlated multi-spiked tensor recovery, including phase-transition analysis, statistical guarantees, and an asymptotically unbiased signal-strength estimator.

Jan 2025 – Present

Lille, France

Independent Researcher

Independent AI/ML Projects

  • Trained and released PetitGPT, a 30-layer, 124.6M-parameter language model, from tokenizer training through approximately 13B pretraining positions on one RTX 4090.
  • Built instruction-tuning and checkpoint-interpolation workflows, evaluated post-training trade-offs, and published native inference, benchmark protocols, training curves, and reproducibility documentation.

Earlier research appointments

  • Inria and École Polytechnique

    Researcher

    Nov 2019 – Oct 2023

  • University of Chicago, Department of Mathematics

    L.E. Dickson Instructor

    Sep 2018 – Aug 2019

  • University of Chicago, Department of Statistics

    Postdoctoral Researcher

    Oct 2016 – Aug 2018

  • CNRS / Université Grenoble Alpes / GIPSA-lab

    Postdoctoral Researcher

    Oct 2013 – Sep 2016

03 / Capabilities

Technical Skills

A focused toolkit for machine learning engineering and mathematically grounded research.

  • Machine Learning

    Deep learning, transformer language models, pretraining, supervised fine-tuning, distillation, DPO, LLM evaluation, experiment design

  • Programming / Frameworks

    Python, PyTorch, NumPy, scikit-learn, TensorFlow

  • Mathematical / Applied Research

    Probability, statistics, optimization, tensor methods, signal processing, optimal control

04 / Research

Selected Research

Selected themes from a broader research record in mathematics, statistics, and algorithms.

  1. 01

    High-dimensional statistics & tensor recovery

    Statistical recovery, phase transitions, and optimization in structured high-dimensional models.

  2. 02

    Tropical methods for data analysis

    Geometric approaches to regression, principal component analysis, and data analysis.

  3. 03

    Optimization & optimal control

    Mathematical and computational methods for optimization and controlled dynamical systems.

For papers, citations, and the complete publication record:

View Google Scholar (opens in a new tab)

05 / Education & contact

Education

  • Ph.D. in Mathematics Texas A&M University, 2013

  • Master's Degree in Mathematics Peking University, 2007

Contact

Email