Role: ML Research Intern (Applied Deep Learning)

Remote or Bengaluru | Part-time or full-time | 2 months, extendable to 6

About TPH
The Product Highway is an AI-native product strategy and engineering firm that partners with businesses from the earliest spark of an idea all the way through to enterprise scale. We grew from zero to multi-million ARR within our first year, working with clients across India, APAC, Europe, and North America.
We only take on problems we find interesting and worth solving, and we refuse to ship anything we wouldn't stand behind. That applies to research the same way it applies to production code.
This internship sits outside our client delivery work. It is a research track: one problem, run properly, over two months. The intended output is publishable.

Who we're looking for
Every person at TPH shares a few traits, regardless of role.
  • End-to-end owners. You own outcomes, not tasks. If something you're responsible for falls through a crack, that's on you.
  • Clear, direct communicators. Bad news doesn't get better with age. You surface problems early and you write with precision.
  • Uncompromising on quality. You have a quality bar that's yours, not your manager's.
  • AI-obsessed. You see AI as how work gets done. If you're doing something manually that AI could handle, you feel that as friction.
  • Structured thinkers. Given a vague problem, you break it down and arrive at a position instead of waiting to be told the answer.
For this role we're not looking for years of experience. We want someone who has already gone past coursework on their own: a paper you reimplemented, a model you fine-tuned on something nobody assigned you, a dataset you built because it didn't exist yet.

Why we're hiring
We have a research problem that needs a dedicated pair of hands for a defined window. It covers the whole pipeline. Source and scrape the data, clean and structure it, build the evaluation setup, then apply and adapt transformer models against it. We want two to three papers out of it, with you as a named author.
Most internships hand you a slice of someone else's work. This one has the opposite problem. You get the whole pipeline, which also means you get every part of it that nobody enjoys. If the data collection is broken, the modelling is worthless, and nobody else is going to fix your scraper for you.
You'll work directly with our leadership and a domain expert. No layers in between.

What you'll do
  • Build the data collection pipeline from scratch: scraping, API ingestion, deduplication, storage. Realistically this is your first two weeks and a chunk of every week after that
  • Clean, label and structure raw data into training and evaluation sets that survive scrutiny
  • Design the evaluation setup before the modelling starts. Baselines, metrics, splits, and what would count as a real result rather than a lucky one
  • Fine-tune and adapt transformer models against the problem. Depending on where it goes, that could mean encoder models, LLM fine-tuning, PEFT/LoRA, or embedding and retrieval work
  • Run experiments reproducibly. Version the data, log the runs, track the configs. An experiment you can't rerun didn't happen
  • Read the surrounding literature and bring back what applies. You should be able to tell us what has already been tried and why it did or didn't work
  • Write up method and results in publishable form, with support from the team on framing
  • Report progress weekly in writing, including what failed. A negative result reported on Friday is useful. The same result surfaced two weeks late is expensive

Must-haves
  • Currently in your 2nd or 3rd year of an undergraduate program (CS, math, stats, or equivalent), or a Master's student with time to commit
  • Available for a continuous two-month block, either full-time or roughly 20 hours a week alongside college. Either works, but tell us which one you're signing up for
  • Strong Python. Pandas, numpy and PyTorch. You should be able to write a training loop without copying one
  • You have scraped something real. requests or httpx, BeautifulSoup or Scrapy, pagination, rate limits, HTML that fights back. Not a tutorial run against a clean site
  • A working understanding of how transformers actually work. Attention, tokenisation, embeddings, what pretraining and fine-tuning do to a model. You should be able to explain why attention is quadratic without looking it up
  • You have fine-tuned or adapted a pretrained model at least once. On your laptop, in Colab, on rented compute, we don't care where. HuggingFace transformers and datasets should be familiar
  • You can read an ML paper and pull the method out of it, including the parts the authors are quietly not telling you
  • Comfortable with git and working in a shared repo

Strong pluses
  • A public GitHub with real projects on it, a Kaggle record, or a paper you have already written or co-written
  • You have built a dataset that didn't exist before, especially from messy or non-standard sources
  • LoRA, QLoRA or other PEFT work on constrained compute
  • Experiment tracking with Weights and Biases, MLflow, or a disciplined homegrown setup
  • Vector databases, embedding retrieval, or RAG pipelines
  • Audio, vision or multimodal data alongside text
  • Any open source ML contribution, however small

AI-native expectations
  • Cursor or Claude Code as your primary development environment. Hand-writing scaffolding, plots and data wrangling is time you don't have
  • AI for literature work. Summarising papers, tracing citations, checking whether your idea has already been tried and published
  • You read and verify everything the model produces. AI-generated code that silently mislabels your training set isn't a bug, it's a wasted month
  • AI on the tedious parts so you have more time on the parts that need judgment, like experiment design and error analysis

Not the right fit if
  • You want to skip the data work and only touch the modelling. Most of this internship is data work, and that is the job
  • You want a structured internship with a curriculum and a daily standup. The problem is defined. The path isn't
  • You can't commit a continuous block. Exams landing in the middle will break the timeline
  • Your ML experience is entirely coursework and tutorial notebooks
  • You expect a paper to be handed to you. Authorship follows contribution

Logistics
  • Duration: 2 months to start. If the work is good, we extend to 6 months
  • Commitment: full-time or part-time at roughly 20 hours a week. Your call, flexible either way
  • Location: remote. If you're in Bengaluru, we'd like you in the office for working sessions
  • Stipend: [confirm before posting]
  • What you get out of it: named authorship on the papers, a completion certificate, and a reference from the team
How to apply
Send your resume, your GitHub and a short note to anju@theproducthighway.com.
In the note, tell us about one ML project you built that nobody assigned you. What you were going for, what broke, and what you'd do differently now. Two paragraphs is plenty. Skip the cover letter formulas, we can tell.