Data & Machine Learning Engineer Intern (Fall semester)
This is for Fall semester 2026!
How to apply
Email hiring@getquikturn.io with your resume and your GitHub profile or a link to something you have built. In a few sentences, tell us about a data or machine learning project you’ve built, why you built it, and how you built it. Please do not send a cover letter.
About Quikturn
Quikturn is the AI-native platform investment bankers, consultants, finance teams, and corporate teams use to build PowerPoint decks. We deliver millions of company logos, deal data, and market and company intelligence through a web app, a native PowerPoint add-in, and an API as a Service (APIaaS).
Our AI deck engine builds complete, draft-ready market maps, sector research reports, pitch decks, and deal materials such as CIPs, CIMs, and teasers in minutes. The intelligence layer behind those outputs covers millions of companies, deals, and investors—replacing much of the manual work analysts spend finding sources, cleaning data, comparing businesses, and structuring a deck.
We are a small founding team moving quickly and closing our first enterprise deals. Interns work on real data and production systems with direct founder context, engineering review, and room to own an important problem deeply.
The role
You will work on the intelligence layer that turns fragmented source material into structured entities, classifications, embeddings, relationships, and retrieval results that Quikturn can use in customer-facing products.
This is applied data and machine learning engineering, not prompt experimentation. A model output is only useful if we know what it was compared against, how often it fails, which failures matter to users, and how the surrounding system detects or contains those failures. You will work across data pipelines, models, evaluation, and product feedback loops to improve that standard.
What you may own and build
- Source ingestion and crawling: Find, fetch, filter, and prioritize the web pages and data sources most likely to contain useful company, deal, investor, ownership, and market information.
- Structured extraction: Turn unstructured text and documents into consistent company profiles, transaction records, investment relationships, and other product-ready fields with traceable provenance.
- Classification and taxonomy: Improve industry, NAICS, software, and market-segment classification across millions of companies, including confidence scoring and behavior for ambiguous businesses.
- Embeddings and semantic retrieval: Generate and evaluate vector representations used for company similarity, natural-language search, market mapping, and candidate discovery.
- Reranking and result quality: Improve the ordering of retrieved companies and evidence so the strongest candidates reach the deck engine instead of being buried in a technically relevant but unhelpful result set.
- Entity resolution: Match companies, domains, brands, investors, subsidiaries, and deals across noisy sources without collapsing distinct entities or duplicating the same one.
- Freshness and drift: Detect when source content, taxonomies, embeddings, or classifications have become stale and build practical paths for reprocessing them.
- Evaluation and labeling: Create representative datasets, annotation guidance, error taxonomies, metrics, and review workflows that show whether a change improved the product.
- Logo intelligence: Contribute to computer-vision and similarity work for logo validation, duplicate detection, background removal, and candidate ranking when it intersects with the broader entity and data-quality problem.
- Large-scale backfills: Build efficient, restartable processing for recalculating classifications or embeddings across millions of rows while preserving reviewed data and making failures observable.
What you will do
- Start with a product failure—an irrelevant company, wrong classification, duplicated entity, stale profile, or unsupported claim—and trace it back through retrieval, models, extraction, crawling, and source data.
- Define success before tuning the system: choose the dataset, baseline, metric, and error categories that make an improvement meaningful.
- Build reproducible training, evaluation, or processing code rather than one-off notebook results.
- Compare model quality, latency, throughput, and cost, then choose the simplest approach that meets the product requirement.
- Design confidence thresholds and human-review paths for cases where automation should not pretend to be certain.
- Work with product and backend systems so improvements reach the web app, PowerPoint add-in, deck engine, and APIaaS—not just an offline experiment.
- Use autonomous developer agents such as Claude Code to accelerate implementation and analysis while personally reviewing the code, methodology, and conclusions.
- Communicate results honestly, including negative results and cases where the available data does not support a confident answer.
Technology
We work primarily with Python, TypeScript, Rust/Go, and SQL, alongside vector search, model-serving, large-scale data pipelines, and applied computer-vision tools. We care less about experience with a particular vendor or model family than disciplined evaluation, sound software fundamentals, and the ability to explain why a result should be trusted.
What we are looking for
- You are a current undergraduate student.
- You are comfortable programming in Python, TypeScript, or another language used for data-intensive work.
- You understand core concepts such as classification, embeddings, train/validation/test separation, precision and recall, and sources of data leakage or bias.
- You can use SQL to inspect data and test a hypothesis.
- You have completed a data, machine learning, information retrieval, computer vision, or applied research project where you measured the result.
- You are curious about why a system failed and patient enough to inspect examples instead of relying only on an aggregate score.
- You write clearly and can distinguish measured evidence from inference.
Helpful, not required
- Experience with web crawling, document extraction, search, vector databases, reranking, or retrieval-augmented systems.
- Familiarity with industry taxonomies, entity resolution, knowledge graphs, or financial datasets.
- Exposure to GPU inference, model serving, batch processing, or large backfills.
- Experience creating labeled datasets, evaluation harnesses, or human-review workflows.
- Experience with image similarity, OCR, logo recognition, or other applied computer-vision problems.
- Experience directing AI coding agents while independently verifying their work.
Internship details
- Type: Semi-paid, temporary/seasonal internship
- Eligibility: Current undergraduate students only. Not open to graduate or master’s-seeking students.
- Location: Remote, or hybrid in New York City
- Schedule: Part-time, 30 hours per week
- Compensation: Intern revenue-split compensation pool
- Work-Study program: No