Work — seven case studies, every number sourced

Results first, then how they happened.

Seven pieces of work, each written the same way: the problem, what I did, what changed. The numbers are the ones I can stand behind. Where the data is private, the source says so.

40.9% of meetings in the morning, against a 30–35% peer target internal data · methodology on request

Course-scheduling study, NYIT

Graduate Research Assistant, Data Science · Jan 2025–Jun 2026

Challenge
The department needed a factual picture of how sections, meeting times, and enrollment were actually distributed across the schedule, and where that diverged from peer targets.
Action
Cleaned and reconciled 15,242 raw registrar records into 8,729 sections and 11,803 meeting instances, then defined and computed the metrics that mattered: time-of-day share, overenrollment by format, Friday load.
Result
40.9% of meetings fell in the morning against a 30–35% peer target; 24.2% of lectures were overenrolled; 10.0% of meetings sat on Fridays. Scope: 3,300+ undergraduates. Findings were presented in department meetings and implemented in scheduling changes for upcoming classes.
MORNING MEETINGS target 30–35% 40.9% OVERENROLLED LECTURES 24.2% FRIDAY MEETINGS 10.0%
Share of class meetings — 15,242 registrar records, 8,729 sections. Bars scaled to a 45% axis.
0.79 ROC-AUC on 6,000 held-out applicants live demo ↗ GitHub ↗

Credit-risk service

Personal MLOps build · Sep 2026 · live on Hugging Face

Challenge
Most model repos stop at a notebook. The goal here was a credit-default model shipped like a product: versioned data, reproducible metrics, and a service a recruiter can open and score against.
Action
Built a DVC pipeline over the UCI Taiwan credit-card dataset, a tested FastAPI endpoint (211 tests, 88% coverage), and CI that re-runs the pipeline and fails if metrics drift beyond ±0.005 — plus secret scanning, a Docker smoke test, and CD to GHCR and a live Space with drift monitoring.
Result
0.79 ROC-AUC and 0.57 PR-AUC on the held-out split, p95 latency 382 ms at concurrency 10. Every number on the demo page is read from a file the pipeline wrote — nothing is typed by hand.
0.81 accuracy on 2,068 held-out tweets live demo ↗ GitHub ↗

Tweet emotion pipeline

Personal MLOps build · Sep 2026 · live on Hugging Face

Challenge
An NLP result you can't reproduce isn't evidence. The aim was a text classifier where every artifact is byte-stable across machines and every public number has a file behind it.
Action
Built a six-stage DVC pipeline over crowd-labeled tweets, every stage driven by params.yaml, compared logistic regression, naive Bayes, and XGBoost by 5-fold cross-validation, and shipped the winner behind a tested FastAPI service (500+ tests, 96% coverage) with CI/CD to GHCR and a live Space.
Result
0.81 accuracy and 0.889 ROC-AUC on the held-out split. One deliberate call: negation words stay out of the stop list, so "not happy" no longer scores like "happy".
40+ AI/ML and full-stack products shipped to clients Upwork profile ↗

Production AI systems for clients

Independent · Apr 2025–now · New York

Challenge
Clients want LLM features that behave in production, not in a demo: predictable outputs, human review where the stakes are high, and evidence that the system works before it ships.
Action
Designed and shipped the systems end to end: RAG pipelines, LLM agents with human-in-the-loop review, tool calling and structured outputs, evaluation harnesses, and the full-stack products around them (Python, React, Node.js, AWS, Docker, CI/CD).
Result
40+ products shipped. Two anonymized examples: a SaaS product delivered with a 98-test suite, and a system processing 600+ emails per weekday across 45 mailboxes. Client names stay private.
85% résumé–job matching accuracy live demo ↗ GitHub ↗

JobCraft

Personal project · Jun–Sep 2025 · open source · live on Hugging Face

Challenge
Keyword search misses the real overlap between a résumé and a role, so job seekers wade through listings that were never a fit.
Action
Built a multi-agent system that reads the résumé and the posting and scores fit through semantic matching over vector embeddings with NLP scoring (Python, LangChain, OpenAI API), then containerized it on AWS EC2 with Docker and GitHub Actions CI/CD.
Result
85% résumé–job matching accuracy. The code has been public on GitHub since June 2026; it now runs as a public demo on Hugging Face — upload a résumé and watch it match.
94% query accuracy on the evaluation set live demo ↗ GitHub ↗

MedBot

Personal project · Feb–May 2025 · live on Hugging Face

Challenge
A medical Q&A assistant that guesses is worse than none. Answers had to be grounded and measurable.
Action
Built a retrieval-grounded (RAG) medical question-answering assistant: trusted medical PDFs embedded into a Pinecone vector knowledge base with Sentence Transformers, orchestrated with LangChain, and served through a Flask app.
Result
94% query accuracy on the evaluation set, with precision, recall, and F1 tracked alongside. Now running as a public demo on Hugging Face — ask it a medical question and it answers only from the reference, or says it does not know.
+15% revenue, with reporting time cut by 30% internal · details on request

Operations analytics, NSK

Data & operations analytics · part-time · Aug 2024–now · petroleum wholesale, Brooklyn

Challenge
Pricing, sales, and day-to-day reporting were manual, so trends and problems surfaced late.
Action
Built the pricing and sales analysis layer over operations data (Python, Excel), automated the recurring reports, and turned the findings into KPIs the owners now track.
Result
Pricing and sales trends contributed to a 15% revenue increase; automated reporting cut manual processing time by 30%. Ongoing.

More builds

Classification

Disease prediction

Supervised classification build predicting disease.

GitHub ↗

Computer vision

Handwritten character recognition

Recognition over handwriting samples.

GitHub ↗

Analysis

Behavioral risk factors

Risk-factor analysis across tobacco-use data.

GitHub ↗

Portfolio

Data science mastery

One structured repo, exploration through prediction.

GitHub ↗

Stack

What I work with

Analysis & modeling

  • Python · SQL
  • pandas · NumPy
  • scikit-learn · statistics
  • PyTorch · TensorFlow
  • CNNs · computer vision
  • NLP · Excel

LLM systems

  • RAG · LangChain · LangGraph
  • OpenAI API · Hugging Face
  • Pinecone · FAISS · vector search
  • Sentence Transformers
  • WhisperX · transcription
  • Tool calling · evaluation harnesses

Shipping

  • AWS · Docker
  • Git · GitHub Actions CI/CD
  • DVC · MLflow
  • FastAPI · Flask · Streamlit
  • React · Node.js

Contact

Building something with data? Say hi.

A role, a collaboration, or a question about anything on this page — my inbox is open, and email is the fastest way in. No deck required.

zulqarnainhsyed@gmail.com