I'm an engineer who works on making machine learning systems measurably trustworthy. At Virginia Tech I test whether fraud-detection methods for a GSMA-partnered SIM-farm project hold up on new data, using leakage checks, calibration and time-based validation. My paper at the MATH-AI workshop asks what a failed proof check really means when a language model proposes a mathematical conjecture. Before graduate school I built personalization models and data pipelines in industry, and I'm now building an LLM inference server from scratch.
I like problems where the first answer is easy and the honest answer is hard: a model that scores well on the wrong split, a proof that fails for the wrong reason. I write the check before I trust the result.
Model evaluation
Fraud and anomaly detection
LLM inference systems
AI for mathematics
Time-series imputation
Experimentation
Data engineering
A small demonstration
A curve is rebuilt from a handful of samples of a signal. It is the same sparse-to-dense problem as my cellular signal work. Slide to change how many samples it gets.
Error on the points it was given
0.000
Error on points it never saw
0.000
Grey line: the true signal. Blue line: the rebuilt curve. Dots: the samples it was given. The curve passes through every dot, so it scores perfectly on them. Only the second number says how good the rebuild is. Judging a model on the data it was given is the first mistake I check for.
Highlights
Lifted customer retention 25% and conversion 18% by building personalization and recommendation models at Sprect.com.
Cut client teams' manual data work by about 30% with automated ETL, validation and de-duplication at Novvum.
2nd of 1,061 teams in the SAIR Mathematics Distillation Challenge.
Co-authored a poster accepted at the MATH-AI Workshop at NeurIPS 2026.
Selected work
SIM-farm fraud detection
Virginia Tech with GSMA, Aug 2026 to present
Problem
Detect SIM farms from high-frequency network telemetry, and know whether the detection methods still work on data they haven't seen.
What I built
Python workflows for anomaly detection, automated validation and statistical evaluation, plus checks for distribution shift, temporal autocorrelation, feature importance, leakage, calibration and time-based validation.
Status
Ongoing research under Yaling Yang.
Beyond the Compile Check
Poster, MATH-AI Workshop at NeurIPS 2026, with Yaling Yang
Problem
When a language model's math conjecture fails to compile in Lean 4, is it false or just unproven?
What I built
A pipeline that has a model propose conjectures, checks them in Lean 4, and tests the failures to tell false claims from unproven ones. It was run on 240 AI-generated claims.
Understand how language models are served efficiently by building the serving layer myself.
What I'm building
A server written from scratch: one request at a time first, then continuous batching and KV-cache management, benchmarked against vLLM.
Status
In progress.
Cellular signal reconstruction
Research, early stage
Problem
Cellular measurements arrive every few seconds. Can they be rebuilt about ten times finer in time?
What I'm building
Time-series imputation baselines, evaluated only on sessions the model has never seen, so the results can't leak from training data.
Status
Data inspected and evaluation rules set. Training has not started.
Experience
Graduate Research Assistant, Virginia Tech
Aug 2026 to present
Research on fraud detection for a GSMA-partnered project, focused on whether the methods are reliable.
Analyze high-frequency telemetry and build Python workflows for anomaly detection, automated validation and statistical evaluation.
Test generalization with distribution similarity tests, temporal autocorrelation analysis, feature-importance evaluation, leakage checks, calibration analysis and time-based validation.
Software Engineer, founding team, Sprect.com
Jan 2024 to May 2025
Built the models and experiments behind customer personalization at an early-stage company.
Built personalization and recommendation models from behavioral and product data, raising customer retention by 25% and conversion by 18%.
Designed and evaluated A/B tests: hypotheses, primary metrics, guardrails, sample-size considerations and confidence intervals, ending in decision-ready recommendations.
Ran funnel, cohort, segmentation, retention and campaign analyses to find behavioral drivers and size opportunities.
Built and validated propensity, churn-risk and customer-value models in Python and SQL, with leakage checks, calibration and drift monitoring.
Data Engineer, Novvum
Jan 2022 to Dec 2023
Turned SaaS clients' raw data into reporting their teams could rely on.
Built SQL ETL pipelines, REST API integrations, Power BI dashboards and reporting workflows for customer, product and operational insight.
Automated transformation, validation, de-duplication and quality control, cutting client teams' manual effort by about 30%.
Developed measurement frameworks for customer engagement, operational performance and business outcomes.
Publications and recognition
Devanshu Dixit, Yaling Yang. Beyond the Compile Check: Separating False from Unproven in LLM Conjecture Generation. Poster, MATH-AI Workshop at NeurIPS 2026.
2nd of 1,061 teams, SAIR Mathematics Distillation Challenge, where I built AI approaches for mathematical reasoning.
Education
M.S. Computer Science, Virginia Tech
Expected May 2027
B.S. Computer Engineering, University of Mumbai
2022
Tools I use most: Python, SQL, R, Pandas, NumPy, Spark, Databricks, Power BI.
Beyond work
Away from the screen I lift and swim. On it, I'm usually chasing a research question I can't leave alone.