MBAI, homeBook a first session
Menu

Interactive tool

Which AI model should you actually use?

Answer two questions and get the right kind of model for your problem, in plain terms and in technical detail. Includes a map of 145 approaches across 14 families, because language models are one of them.

The mistake is treating "AI" and "chatbot" as the same word. A large language model is a remarkable reasoner over language. It is not a forecaster, not an optimizer, not a purpose-built perception system, and not the economical way to sort things at volume. The pattern that actually wins on cost, speed, and accuracy is quieter: specialist models do the work, and a language model takes the order, routes it, and reports back in plain English.

So tell the tool what you are starting with and what you need back, and it will point you to the kind of model that fits, in plain terms and in technical detail.

Two questions.

Start with what you have and what you need out of it. You get a plain-English answer and a real example, then the technical pick and the trap that catches most people.

Every path in the picker, written out (30 of them)
A spreadsheet or table of records5 answers

Starting with a spreadsheet or table of records, you want to: Predict a number or a likelihood: revenue, risk, who will churn

You want the model to look at each row and predict a value, like a dollar amount or the odds that something happens.

Use this

Gradient-boosted trees: XGBoost, LightGBM, or CatBoost

Boosted trees still beat deep learning on tabular data. They are fast, cheap, give you calibrated probabilities, and explain themselves through SHAP.

In practice

A subscription business scores every account each night for cancellation risk, so the team can call the ten most likely to leave before they do.

The trap

Pasting the row into an LLM prompt. It cannot calibrate a probability, and it will produce a confident number with nothing behind it.

Starting with a spreadsheet or table of records, you want to: A decision you have to explain and defend: lending, insurance, compliance

You need the prediction and a clear reason for it, because a regulator, an auditor, or the customer will ask why.

Use this

A generalized additive model (GAM or EBM), or plain logistic regression

You get close to boosted-tree accuracy while every feature's effect stays a curve a regulator, clinician, or underwriter can look at and argue with.

In practice

A lender approves or declines applicants and can show, factor by factor, what moved each decision, so it holds up in an audit.

The trap

Shipping a black box into a regulated decision, then trying to bolt an explanation on afterward.

Starting with a spreadsheet or table of records, you want to: Find natural groups or segments you did not know were there

You want the data sorted into groups it forms on its own, without you deciding the categories in advance.

Use this

HDBSCAN, or a Gaussian mixture model

Neither one makes you guess the number of clusters up front, and both cope with clusters that are not neat spheres.

In practice

A retailer discovers its customers actually fall into five distinct buying patterns, and builds a different offer for each.

The trap

Forcing k-means when you have no idea what k is. It will happily return exactly the number of segments you asked for, whether or not they exist.

Starting with a spreadsheet or table of records, you want to: Flag the unusual or suspicious ones, with no examples to train on

You want the odd ones out flagged automatically, even though you have never labelled what a bad one looks like.

Use this

Isolation Forest, or an autoencoder scored on reconstruction error

Both learn what normal looks like and flag departures from it, so you need no examples of the thing you are hunting.

In practice

A payments company flags transactions that do not fit a customer's normal pattern, catching fraud it has never seen before.

The trap

Waiting for a labeled fraud dataset. It never arrives, because by the time fraud is labeled it has already cost you.

Starting with a spreadsheet or table of records, you want to: Find who an offer will sway, not who would have bought anyway

You want to spend your discount only on the people it actually changes, not on the ones who were going to buy regardless.

Use this

Uplift modeling, also called CATE: causal forest, or a T-learner or X-learner

This estimates the effect of the offer on each individual, so you spend only on the people the offer actually moves.

In practice

An online store sends the coupon only to shoppers it will genuinely convert, and stops giving away margin to loyal customers who needed no nudge.

The trap

Using a propensity or churn model. It finds who will buy or who will leave. It does not find who you can change, so you end up discounting people who would have bought anyway.

Text or documents5 answers

Starting with text or documents, you want to: Sort a huge pile of text into categories: tickets, emails, reviews

You have far more text than anyone can read, and you want each piece dropped into the right bucket automatically.

Use this

A fine-tuned text classifier: DeBERTa or DistilBERT, or fastText for a cheap baseline

Once fine-tuned on your labels it is orders of magnitude cheaper and faster, and usually more accurate than an LLM doing the same job. Before you have those labels, a zero-shot LLM is a reasonable way to start.

In practice

A support team auto-routes every incoming ticket to the right queue in milliseconds, for a fraction of a cent each.

The trap

One LLM call per document. It works in the demo and the bill becomes indefensible at volume.

Starting with text or documents, you want to: Find the right document, passage, or answer out of many

You want to search a big collection by meaning, not just keywords, and get the most relevant piece back.

Use this

Hybrid retrieval: BM25 plus embeddings, then a reranker

BM25 catches exact strings, embeddings catch meaning, and the reranker sharpens the top few. Together they beat either one alone.

In practice

An internal help desk searches ten years of policies and returns the exact paragraph that answers an employee's question.

The trap

Embeddings by themselves. They quietly fail on part numbers, SKUs, names, and error codes, which is often exactly what people search for.

Starting with text or documents, you want to: Pull specific facts out of documents: names, amounts, dates

You want particular fields lifted cleanly out of documents, with nothing invented.

Use this

Named entity recognition (spaCy, GLiNER) or extractive question answering

Extractive models return the exact span from the source, so they cannot invent a value that was never in the document.

In practice

A finance team pulls vendor, date, and total off thousands of invoices, and never risks a made-up number.

The trap

Asking an LLM to return JSON and trusting it. It will hallucinate a plausible invoice total on the documents where the number is hard to read.

Starting with text or documents, you want to: Find the themes across thousands of comments or reviews

You want to know what people are actually saying at scale, grouped into the themes that genuinely recur.

Use this

BERTopic

It embeds, clusters, and labels the clusters, so the themes come from all of the data rather than from the fifty responses a human had time to read.

In practice

A product team turns 20,000 open-ended survey answers into the eight themes that actually move the score.

The trap

Reading a sample and calling it a theme. You find what you expected to find.

Starting with text or documents, you want to: Write something new, or answer an open-ended question

You genuinely need language written or reasoned through, where the answer is open-ended. This is the language model's real job.

Use this

A large language model. This one is genuinely its job.

Open-ended language under ambiguity is the thing LLMs are unmatched at. Drafting, reasoning, summarizing, and explaining are home turf.

In practice

A marketing team drafts first-pass copy, and a support agent gets a suggested reply it can edit and send.

The trap

Assuming that because it is right here, it is right everywhere else on this list.

Images or video6 answers

Starting with images or video, you want to: Sort images into categories: pass or fail, product type

You want each image given a single verdict or category.

Use this

A CNN (EfficientNet or ConvNeXt), or a vision transformer

Small, fast, and runs on the edge. A CNN will do this on a device with no network connection and no per-call cost.

In practice

A manufacturer tags each photo off the line as pass or fail, on a device with no internet and no per-image cost.

The trap

Sending every frame to a multimodal LLM API and paying for it forever.

Starting with images or video, you want to: Find and count objects in an image or video

You want to know what is in the frame, where it is, and how many.

Use this

Object detection: YOLO or RT-DETR

You get boxes, classes, and counts in milliseconds, which is what you need to audit a shelf, check PPE, or count vehicles.

In practice

A warehouse counts pallets and checks that every worker is wearing a helmet, straight from the cameras already on the wall.

The trap

Asking a vision LLM to count things. It is unreliable at counting and will not give you coordinates.

Starting with images or video, you want to: Measure the exact area or outline of something

You need the precise shape or area of something in the image, not just that it is present.

Use this

Segmentation: U-Net, or SAM 2

Pixel-level boundaries are what let you measure a tumor volume or the surface area of a defect, rather than just noting it exists.

In practice

An insurer measures the exact area of roof damage from an aerial photo, and sizes the claim from it.

The trap

Settling for a bounding box when the number you actually need is an area.

Starting with images or video, you want to: Read the text inside a scan or photo: invoices, forms, IDs

You want the words trapped inside an image turned into clean, structured data.

Use this

OCR plus a layout model: PaddleOCR or Surya, then LayoutLMv3 or Donut

OCR is a pipeline, not a model. Detect the text, recognize the characters, understand the layout, normalize the fields, and only then let an LLM clean up the ambiguous tail.

In practice

An operations team turns scanned invoices into structured records accurately, across hundreds of thousands of pages.

The trap

Pointing a multimodal LLM at the raw scan. It works on twenty documents and falls apart at two hundred thousand, on both accuracy and cost.

Starting with images or video, you want to: Spot defects on a line when you only have good samples

You want faults flagged even though you only have pictures of good parts, not every possible defect.

Use this

PatchCore or PaDiM

They learn what a good part looks like and flag anything that departs from it, so you never have to collect examples of every possible defect.

In practice

A factory catches scratches and misprints in real time, without ever assembling a library of every way a part can fail.

The trap

Trying to build a labeled dataset of every failure mode. The rare defects are the expensive ones and you will never have enough of them.

Starting with images or video, you want to: Create a new image from a description

You want an image made to order, that you can keep on brand and reuse.

Use this

Diffusion: Stable Diffusion, Flux, or a hosted equivalent

Diffusion gives the highest fidelity, and ControlNet or LoRA let you hold pose, composition, and brand style steady.

In practice

A creative team generates on-brand product concepts and marketing visuals, controlled to match house style.

The trap

Generating without conditioning, then wondering why nothing is on brand or reusable.

Audio or speech4 answers

Starting with audio or speech, you want to: Turn a recording into text: calls, meetings, voicemails

You want spoken audio converted into an accurate transcript.

Use this

Whisper, or Parakeet

Purpose-built speech recognition, far cheaper than a multimodal model, and it runs locally if the audio is sensitive.

In practice

A sales team gets every call transcribed automatically, kept on their own systems when the content is sensitive.

The trap

Treating transcription and understanding as one step. Transcribe first, then let an LLM summarize the text.

Starting with audio or speech, you want to: Tell who spoke and when in a multi-person recording

You want the transcript split by speaker, so you know who said what.

Use this

pyannote for speaker diarization

Diarization is a separate task from transcription. Chain it with Whisper to get an attributed transcript.

In practice

A firm produces interview and multi-party call transcripts that are labelled by speaker, not one undifferentiated block.

The trap

Expecting the transcription model to also tell you who was talking. It will not.

Starting with audio or speech, you want to: Recognize a sound that is not speech: a machine fault, an alarm

You want a non-speech sound identified: a warning, a fault, an event.

Use this

Audio event classification: YAMNet, PANNs, or AST

The signal is a sound, not a word: a failing bearing, breaking glass, an alarm. Speech models are the wrong instrument entirely.

In practice

A plant listens for the specific sound of a bearing starting to fail, and raises an alert before it breaks.

The trap

Transcribing audio that contains no speech and wondering why the output is empty.

Starting with audio or speech, you want to: Turn text into a natural spoken voice

You want written text read aloud in a natural voice.

Use this

Text to speech: ElevenLabs, Kokoro, or Piper

Modern TTS is natural enough for production voice agents, IVR, and accessibility, and the open options run on your own hardware.

In practice

A company powers its phone menu and accessibility features with natural speech, on hardware it controls.

The trap

Cloning a voice without written consent. The technical bar is low and the legal bar is not.

Numbers that move over time4 answers

Starting with numbers that move over time, you want to: Forecast what comes next: demand, sales, load

You want a forecast of future values, with an honest sense of how confident it is.

Use this

Start with ETS or Prophet. Escalate to LightGBM on lagged features, then a Temporal Fusion Transformer.

Climb the ladder only as far as the accuracy actually requires. Boosted trees on lags win most real forecasting problems and are cheap to run.

In practice

A retailer forecasts demand for every product in every store, orders the right stock, and ties up less cash doing it.

The trap

Handing a spreadsheet to an LLM and asking it to forecast. It cannot do arithmetic reliably over a long series and has no concept of seasonality or a confidence interval.

Starting with numbers that move over time, you want to: Predict when something will happen, not just whether

You want the timing of an event, not just a yes or no.

Use this

Survival analysis: Cox proportional hazards, or DeepSurv

Survival models handle censored data, which is the whole difficulty: most of your customers have not churned yet, and that is information, not a gap.

In practice

A subscription business predicts when each customer is likely to cancel, so it can step in during the right window.

The trap

Turning a timing question into a yes or no classification and throwing away the timing.

Starting with numbers that move over time, you want to: See the range of outcomes, not just one number

You want the spread of what could happen, so you can plan for the downside, not just the average.

Use this

Monte Carlo simulation, or a probabilistic forecaster such as DeepAR

A point estimate hides the thing you are actually deciding on, which is the spread of outcomes.

In practice

A finance team shows the board a realistic best-to-worst range for next quarter's cash, not a single misleading number.

The trap

Reporting one number to a board that is really asking about the downside.

Starting with numbers that move over time, you want to: Detect when a trend genuinely shifted or broke

You want to know when a real change happened, separate from the normal ups and downs.

Use this

Change-point detection (PELT, BOCPD), or a streaming anomaly model

It tells you a structural break happened and when, rather than leaving you to eyeball a chart.

In practice

An analytics team is alerted the day conversion actually dropped, instead of arguing over a noisy chart a week later.

The trap

Declaring a change from a dashboard by eye. Seasonality fools everybody.

A network of connected things2 answers

Starting with a network of connected things, you want to: Flag suspicious rings or links between people, accounts, or transactions

You want to catch patterns that only show up in how things connect: rings, chains, shared links. You never have to hear the word node.

Use this

A graph neural network: GraphSAGE or GAT

The connections are the signal. Fraud rings, laundering chains, and supply-chain exposure are invisible when each row is scored on its own.

In practice

A bank detects fraud rings and laundering chains that look perfectly normal one account at a time.

The trap

Flattening the graph into a table of features. You throw away the exact structure you were trying to detect.

Starting with a network of connected things, you want to: Work out which records are really the same person or company

You want to know when two records are actually the same real-world person or business, entered or spelled differently.

Use this

Entity resolution with embeddings plus a graph, or a Siamese network

Similarity plus connectivity resolves duplicates that neither signal catches alone.

In practice

A company merges duplicate customer records and stops treating one client as three different accounts.

The trap

Fuzzy string matching by itself. It merges two real companies with similar names and splits one company with two spellings.

Customers and how they behave3 answers

Starting with customers and how they behave, you want to: Recommend what a customer should see or buy next

You want to suggest the next thing for each person, learned from how everyone actually behaves.

Use this

Two-tower retrieval to shortlist, then a ranking model such as DeepFM

Retrieve a thousand candidates from millions in milliseconds, then spend real compute ranking only those.

In practice

A marketplace shows each shopper the products they are most likely to want next, chosen from millions, instantly.

The trap

Asking an LLM for recommendations. It has never seen your behavioral log, which is the only thing that actually predicts the next click.

Starting with customers and how they behave, you want to: Find the winning option without wasting traffic on losers

You want to test options and automatically shift toward the winner, instead of splitting traffic evenly for weeks.

Use this

A multi-armed or contextual bandit: Thompson sampling

A bandit shifts traffic toward the winner while it is still learning, so you stop paying to show people the losing variant.

In practice

A site tests five homepage layouts and quietly sends more visitors to the best one as the results come in.

The trap

A fixed-split A/B test that runs for six weeks while half your traffic sees the worse option.

Starting with customers and how they behave, you want to: Find the best plan under hard rules: schedules, routes, budgets

You want the best possible plan that never breaks a hard rule, like a shift that cannot be double-booked.

Use this

Mixed-integer programming or a constraint solver: OR-Tools, Gurobi

Scheduling, routing, and allocation have constraints that cannot be violated, and a solver gives you a provably optimal, feasible plan.

In practice

A field-service company schedules 400 technicians against 3,000 jobs, with no impossible or double-booked assignments.

The trap

An LLM-generated schedule. It will look plausible, read well, and quietly break three hard constraints.

Several of these at once1 answer

Starting with several of these at once, you want to: See how the pieces fit together into one system

You have more than one of these needs, and you want to know how the parts combine into something that works.

Use this

Specialist models underneath, one LLM out front

This is the architecture that wins on cost, latency, and accuracy at the same time. The LLM reads the request, calls the forecaster, the OCR, and the retriever, then writes the result up in plain English. Each specialist does what it is good at, and the language model does the language.

In practice

A claims process reads the documents with OCR, estimates the cost with a forecaster, flags the odd ones with an anomaly model, and an LLM writes the summary and routes it. Each part does what it is best at.

The trap

Making the LLM do the work instead of the routing. That is how you end up paying frontier prices to do logistic regression, badly.

The short version

Twelve jobs where a language model is the wrong tool.

Every row is something people currently hand to a chatbot. Every row has a better answer that is cheaper and faster, and once you have labelled data, usually more accurate.

Twelve jobs people hand to a language model, and what fits each one better.
The jobUsually reached forWhat actually fitsWhy
Predict churn or default from customer attributesAn LLM with the row in the promptGradient-boosted treesBoosted trees beat deep learning on tabular data. Cheaper, faster, calibrated, and explainable through SHAP.
Forecast next quarter's demandAn LLM asked to analyze the dataETS, Prophet, or LightGBM on lagsLLMs cannot do arithmetic reliably over a long series, and have no notion of seasonality or a confidence interval.
Sort ten million support tickets into twelve bucketsOne LLM call per ticketA fine-tuned DeBERTa or fastTextOrders of magnitude cheaper and faster, and usually more accurate once trained on your own labels.
Find the documents similar to this oneAn LLM reading everythingEmbeddings, BM25, and a rerankerRetrieval is a search problem, not a generation problem.
Spot a defect on the production lineA multimodal LLMPatchCore, PaDiM, or a small CNNRuns on the device in milliseconds. No round trip, no token cost, no network dependency.
Read a scanned invoiceA vision LLM on the raw imageOCR and a layout model, LLM for the tailPurpose-built OCR is more accurate on dense text and far cheaper. Use the LLM to clean up, not to extract.
Choose who gets the discountLLM judgementUplift or CATE modelingYou do not want who will buy. You want who will buy only if you discount. Different question, different math.
Schedule 400 technicians across 3,000 jobsAn LLMA constraint solver or MILPCombinatorial optimization with hard constraints. An LLM returns a plausible, infeasible schedule.
Recommend the next productAn LLMTwo-tower retrieval and a rankerRecommenders learn from your behavioral log. The LLM has never seen it.
Decide which of two headlines convertsAn LLM's opinionA contextual bandit or an A/B testGet the answer from reality, not from a prior.
Transcribe a two-hour callAn LLMWhisper, then an LLM to summarizeRight tool per stage. Chain them rather than collapsing them into one.
Catch an anomaly in server metricsAn LLMIsolation Forest or an autoencoderStreaming, unlabeled, and real time. An LLM cannot sit in that loop.

The wider field

Language models are one family out of fourteen.

Every family in this guide, sized by how many approaches it holds.

  • Classical and tabular ML, 12 approaches
  • Unsupervised learning and structure, 14 approaches
  • Time series and forecasting, 13 approaches
  • Neural architectures, 16 approaches
  • Computer vision, 19 approaches
  • Audio and speech, 10 approaches
  • Language models (LLM), 6 approaches
  • Language work without generation, 10 approaches
  • Recommender systems, 7 approaches
  • Reinforcement learning and decisioning, 7 approaches
  • Optimization and operations research, 7 approaches
  • Causal inference, 8 approaches
  • Scientific and domain models, 9 approaches
  • The supporting cast, 7 approaches

Now find yours.

Every problem the guide recognises is in the first table below, and every approach it holds is in the second. The search and the map are a faster way through the same rows.

Every problem the search recognises (47 of them)
Every problem the search recognises, and the answer it gives.
Your problemUse thisWhy, and the trap
Sales lead generationGenerating and working sales leads is really four different jobs, and each one wants a different model. One tool cannot do all four well.A stack: gradient-boosted trees to score leads, uplift or CATE to pick who a touch will actually move, similarity or lookalike search to find more accounts like your best ones, and a language model to draft the outreach.Scoring ranks the leads you already have. Uplift finds who your effort changes rather than who was going to convert anyway. Lookalike finds new prospects that resemble closed-won accounts. The language model writes. Different questions, different math.The trap: Asking one language model to 'generate leads.' It writes fluent outreach but cannot rank your pipeline, has never seen your conversion history, and cannot tell you who is actually worth calling.
Lead scoring and prioritizationYou have more leads than your team can work, and you want them ranked by how likely each is to convert.Gradient-boosted trees: XGBoost, LightGBM, or CatBoost.Boosted trees are the default winner on spreadsheet-shaped data like a lead record. They give calibrated probabilities you can threshold, and SHAP shows why each lead scored the way it did.The trap: Pasting the lead into a language model prompt. It returns a confident score with nothing behind it and cannot calibrate a probability.
Churn and retentionYou want to know which customers are about to leave, and ideally when, so you can step in first.Gradient-boosted trees to score who is at risk, and survival analysis (Cox or DeepSurv) when the timing matters.Boosted trees rank risk from account attributes. Survival models add the piece a yes-or-no model throws away: how long until they go, correctly using the customers who have not churned yet.The trap: Discounting everyone flagged at-risk. Many would have stayed anyway. Pair the risk model with uplift so you only spend on the ones a save offer actually moves.
Upsell and cross-sellYou want to suggest the right next product or plan to each existing customer.A recommender (two-tower retrieval then a ranker) for what to offer, and uplift or CATE to decide who is worth the outreach.The recommender learns from what everyone actually bought next. Uplift keeps you from spending on customers who would have upgraded on their own.The trap: Asking a language model what to upsell. It has never seen your purchase log, which is the only thing that predicts the next buy.
Who to give the offer toYou want to spend a discount only on the people it will actually change, not on the ones who would have bought regardless.Uplift modeling, which estimates each person's CATE: a causal forest, or a T-learner or X-learner.This estimates the effect of the offer on each individual, so you spend only where the offer moves the outcome.The trap: Using a propensity or churn model. It finds who will buy or who will leave, not who you can change, so you discount people who were already going to act.
Customer segmentationYou want your customers sorted into the natural groups they already form, without deciding the categories in advance.HDBSCAN or a Gaussian mixture model.Neither makes you guess the number of groups up front, and both handle groups that are not neat spheres, unlike k-means.The trap: Forcing k-means with a guessed number of clusters. It returns exactly that many groups whether or not they are real.
Find more customers like our bestYou want to find new prospects that resemble the customers who already work out well for you.Turn each account into an embedding, then rank prospects by similarity, optionally sharpened with a small classifier trained on best-versus-rest.Similarity over a rich feature or text embedding surfaces the accounts that look like closed-won, which filtering by industry and size alone misses.The trap: Hand-writing a rules filter on industry and headcount. It is brittle and misses the non-obvious matches a similarity model finds.
Marketing attribution and budgetYou want to know how much each channel actually contributes to sales, so you can move budget to what works.Marketing mix modeling (Robyn or Meridian), plus randomized geo experiments where you can run them.Mix modeling is privacy-safe and top-down, and a geo experiment gives a true causal read on a channel rather than a correlation.The trap: Last-click attribution. It hands all the credit to the final touch and quietly defunds the channels that created the demand.
Product recommendationsYou want to show each person the items they are most likely to want next, chosen from a large catalog.Two-tower retrieval to shortlist, then a ranking model such as DeepFM.Retrieve a thousand candidates from millions in milliseconds, then spend real compute ranking only those.The trap: Asking a language model for recommendations. It has never seen your behavioral log, which is the only thing that predicts the next click.
Test which version winsYou want to try options and shift toward the winner, instead of guessing or splitting traffic evenly for weeks.A multi-armed or contextual bandit (Thompson sampling), or a clean A/B test when you need one rigorous number.A bandit moves traffic to the leader while it is still learning, so you stop paying to show the losing variant. An A/B test is right when you need a single defensible result.The trap: A language model's opinion on which headline is better. Get the answer from real traffic, not from a prior.
Demand and sales forecastingYou want a forecast of future demand or sales, with an honest sense of how confident it is.Start with ETS or Prophet, escalate to LightGBM on lagged features, then a Temporal Fusion Transformer only if the accuracy needs it.Climb the ladder only as far as accuracy requires. Boosted trees on lags win most real forecasting problems and are cheap to run.The trap: Handing a spreadsheet to a language model and asking it to forecast. It cannot do reliable arithmetic over a long series and has no concept of seasonality or a confidence interval.
Sentiment analysisYou want each message, review, or ticket scored as positive, negative, or neutral, at scale.A fine-tuned text classifier: DeBERTa or DistilBERT.Once fine-tuned on your labels it is orders of magnitude cheaper and faster than a language model per message, and more consistent. A zero-shot LLM is a fine way to start before you have labels.The trap: One language model call per message. It works in the demo and the bill is indefensible at volume.
Themes across reviews or surveysYou want to know what people are actually saying at scale, grouped into the themes that genuinely recur.BERTopic.It embeds, clusters, and labels the clusters, so the themes come from all of the data rather than the fifty responses a human had time to read.The trap: Reading a sample and calling it a theme. You find what you expected to find.
Sort and route support ticketsYou have far more tickets, emails, or messages than anyone can read, and you want each one categorized and routed automatically.A fine-tuned text classifier (DeBERTa, DistilBERT, or fastText), with a language model only for the ambiguous tail and for drafting replies.The classifier labels and routes in milliseconds for a fraction of a cent, at accuracy a language model cannot beat once it is trained on your labels.The trap: One language model call per ticket to both classify and answer. Split the job: classify cheaply, generate only where you actually need language.
Answer questions from your documentsYou want people to ask a question in plain English and get an answer grounded in your own documents.Retrieval-augmented generation: hybrid retrieval (BM25 plus embeddings plus a reranker) to find the right passages, then a language model to answer from them.The retrieval half is a search problem, not a generation problem, and it is what keeps the answer grounded. The language model only writes the reply from the passages it was handed.The trap: Putting everything in the prompt and skipping retrieval, or trusting the model's memory. That is how you get a confident answer from a document that does not exist.
Fraud detectionYou want suspicious activity flagged, including patterns you have never labelled and rings that look normal one account at a time.An anomaly model (Isolation Forest or an autoencoder) for the unusual, and a graph neural network for rings and laundering chains.Anomaly models need no examples of fraud, which you rarely have in time. The graph model catches connected fraud that is invisible when each transaction is scored on its own.The trap: Waiting for a clean labelled fraud dataset, or flattening the network into a table and throwing away the exact structure you were trying to detect.
Anomaly and outlier detectionYou want the odd ones out flagged automatically, in data where you have never labelled what a bad one looks like.Isolation Forest or an autoencoder scored on reconstruction error, and change-point detection for streams that shift over time.They learn what normal looks like and flag departures from it, so you need no examples of the thing you are hunting.The trap: Trying to use a language model in a streaming, unlabeled, real-time loop. It cannot sit there, and it has no notion of your normal.
Credit risk and underwritingYou need to predict default risk and be able to explain and defend every decision to a regulator or the customer.A generalized additive model (GAM or EBM), or logistic regression.You get close to boosted-tree accuracy while every factor's effect stays a curve an auditor, underwriter, or regulator can inspect and argue with.The trap: Shipping a black-box model into a regulated decision, then trying to bolt an explanation on afterward.
Anti-money-launderingYou want to catch laundering that only shows up in how accounts and transactions connect.A graph neural network (GraphSAGE or GAT) over the transaction graph.Laundering chains are invisible when each account is scored alone. The connections are the signal.The trap: Flattening the graph into per-account features. You discard the structure that was the whole tell.
Read invoices, forms, and scansYou want the words and fields trapped inside scans or photos turned into clean, structured data, at volume.An OCR pipeline: PaddleOCR or Surya to read, then LayoutLMv3 or Donut to understand the layout, with a language model only for the ambiguous tail.OCR is a pipeline, not a model. Purpose-built OCR is more accurate on dense text and far cheaper than a vision model, across hundreds of thousands of pages.The trap: Pointing a multimodal language model at the raw scan. It works on twenty documents and falls apart at two hundred thousand, on both accuracy and cost.
Search a big pile of documentsYou want to search a large collection by meaning, not just keywords, and get the most relevant passage back.Hybrid retrieval: BM25 plus embeddings, then a reranker.BM25 catches exact strings, embeddings catch meaning, and the reranker sharpens the top few. Together they beat either alone.The trap: Embeddings by themselves. They quietly fail on part numbers, SKUs, names, and error codes, which is often exactly what people search for.
Pull specific facts out of textYou want particular fields lifted cleanly out of documents, with nothing invented.Named entity recognition (spaCy or GLiNER) or extractive question answering.Extractive models return the exact span from the source, so they cannot invent a value that was never in the document.The trap: Asking a language model to return JSON and trusting it. It will hallucinate a plausible value on the documents where the real one is hard to read.
Work out which records are the sameYou want to know when two records are actually the same real-world person or company, entered or spelled differently.Entity resolution with embeddings plus a graph, or a Siamese network.Similarity plus connectivity resolves duplicates that neither signal catches alone.The trap: Fuzzy string matching by itself. It merges two real companies with similar names and splits one company with two spellings.
Financial and cash forecastingYou want a forward view of cash or revenue, and the board is really asking about the range, not one number.A forecaster (ETS, Prophet, or LightGBM on lags) for the central path, and Monte Carlo simulation or a probabilistic model like DeepAR for the range.A point estimate hides the thing you are deciding on, which is the spread of outcomes. Show the best-to-worst band, not a single misleading line.The trap: Reporting a single forecast to a board that is really asking about the downside.
Scheduling and rosteringYou want the best possible schedule that never breaks a hard rule, like a shift that cannot be double-booked.Mixed-integer programming or a constraint solver: OR-Tools, CP-SAT, or Gurobi.Scheduling has constraints that cannot be violated, and a solver returns a provably feasible, optimal plan.The trap: A language-model-generated schedule. It looks plausible, reads well, and quietly breaks three hard constraints.
Route and delivery optimizationYou want the most efficient routes under real limits: time windows, vehicle capacity, driver hours.A vehicle routing solver, or mixed-integer programming (OR-Tools).Routing under hard constraints is combinatorial optimization, where a solver gives you a feasible, near-optimal plan you can trust.The trap: A language model asked to plan the route. It returns something plausible that violates the constraints you cannot violate.
Inventory and reorderingYou want to hold enough stock to meet demand without tying up cash in shelves.A demand forecast (ETS, Prophet, or LightGBM on lags) feeding an optimization or inventory policy.Forecast demand and its uncertainty first, then let an optimization set reorder points against service-level and budget constraints.The trap: Setting one reorder rule by gut across every product. Demand varies too much for a single static number.
Predictive maintenanceYou want a warning before a machine fails, from whatever signal it gives off first.Survival analysis for time-to-failure, anomaly detection on sensor streams, and audio event classification when the tell is a sound.Different machines warn you differently: a drifting sensor, a rising vibration, a bearing that starts to whine. Match the model to the signal, not to a chatbot.The trap: Waiting to collect examples of every failure mode. The rare failures are the expensive ones and you never have enough, so learn what normal looks like and flag departures.
Visual quality controlYou want defects flagged on a line, even though you only have pictures of good parts, not every possible fault.PatchCore or PaDiM for visual anomaly detection, or a small CNN when you can label pass and fail.They learn what a good part looks like and flag anything that departs from it, on the device, in milliseconds, with no per-image cost.The trap: Sending every frame to a multimodal language model API and paying for it forever, at latency a line cannot tolerate.
Did this change actually work?You rolled something out and want the real effect, separated from everything else that was moving at the same time.A randomized experiment where you can run one, otherwise difference-in-differences, synthetic control, or structural causal models (DoWhy, EconML).These isolate cause from coincidence, which a before-and-after chart cannot. That difference is worth a great deal of money.The trap: Reading a dashboard and declaring victory. Seasonality and everything else you changed will happily take the credit.
Classify text at scaleYou want to drop a high volume of text into predefined categories, cheaply and consistently.A fine-tuned text classifier: DeBERTa, DistilBERT, or fastText.These understand text without generating it, at a fraction of a language model's cost, and beat it on accuracy once trained on your labels. Before you have labels, a zero-shot LLM is a reasonable start.The trap: One language model call per document. Cheap in the demo, indefensible at volume.
Translation at volumeYou want to translate a large amount of text between languages.A dedicated machine translation model: NLLB, MarianMT, or SeamlessM4T.Purpose-built translation is a fraction of the cost of a language model per word, which becomes ruinous at bulk volume.The trap: Paying a frontier language model per word to translate a whole corpus.
Summarize, draft, or reasonYou genuinely need language written or condensed, where the answer is open-ended. This is the language model's real job.A large language model.Open-ended language under ambiguity is the thing language models are unmatched at. Drafting, summarizing, reasoning, and explaining are home turf.The trap: Assuming that because it is right here, it is right for the forecasting, scoring, and extraction jobs elsewhere on your list.
Transcribe speech to textYou want spoken audio, like calls or meetings, converted into an accurate transcript, and often split by speaker.Whisper or Parakeet, chained with pyannote when you also need to know who spoke.Purpose-built speech recognition is far cheaper than a multimodal model and runs locally when the audio is sensitive. Diarization is a separate step that labels the speakers.The trap: Treating transcription and understanding as one step. Transcribe first, then let a language model summarize the text.
Text to a natural voiceYou want written text read aloud in a natural voice, for a phone system, narration, or accessibility.Text to speech: ElevenLabs, Kokoro, or Piper.Modern TTS is natural enough for production, and the open options run on your own hardware.The trap: Cloning a voice without written consent. The technical bar is low and the legal bar is not.
Classify or sort imagesYou want each image given a single verdict or category.A CNN (EfficientNet or ConvNeXt) or a vision transformer.Small, fast, and runs on the edge, with no network connection and no per-call cost.The trap: Sending every frame to a multimodal language model API and paying for it forever.
Find and count objectsYou want to know what is in the frame, where it is, and how many.Object detection: YOLO or RT-DETR.You get boxes, classes, and counts in milliseconds, which is what you need to audit a shelf, check PPE, or count vehicles.The trap: Asking a vision language model to count things. It is unreliable at counting and will not give you coordinates.
Generate imagesYou want an image made to order that you can keep on brand and reuse.Diffusion (Stable Diffusion or Flux), steered with ControlNet or LoRA.Diffusion gives the highest fidelity, and conditioning holds pose, composition, and brand style steady.The trap: Generating without conditioning, then wondering why nothing is on brand or reusable.
Generate videoYou want moving footage you cannot practically shoot.A video generation model: Sora, Veo, Runway, or Kling.These create video from a prompt or a still, for concepts and footage that would be expensive or impossible to film.The trap: Expecting frame-perfect control and consistency. Storyboard tightly and expect to curate takes.
Moderate and filter contentYou want unsafe, abusive, or policy-violating content screened out, at the edge of a system open to the public.Safety and moderation classifiers, such as Llama Guard, plus a spam classifier.Small, fast classifiers screen every message cheaply, where a language model per item would be too slow and too costly.The trap: Relying on a prompt in your main model to self-moderate. Screen at the edge with a dedicated classifier.
Find the best plan under hard rulesYou want the best possible plan that never breaks a hard constraint.Mixed-integer programming or constraint programming: OR-Tools, CP-SAT, or Gurobi.Allocation, assignment, scheduling, and routing all have rules that cannot be violated, and a solver returns a provably optimal, feasible answer.The trap: A language model asked to 'optimize' it. It returns a plausible plan that quietly breaks the constraints.
Model what could happenYou want the range of outcomes and the odds, so you can plan for the downside, not just the average.Monte Carlo simulation, or discrete-event and agent-based simulation for a whole system.Simulation shows the spread and the tail, and lets you test a change before it is real.The trap: Reporting one expected number when the decision is really about the risk.
Predict when, not just whetherYou want the timing of an event, not just a yes or no.Survival analysis: Cox proportional hazards, or DeepSurv.Survival models handle the customers or machines that have not had the event yet, which is information, not a gap.The trap: Turning a timing question into a yes-or-no classification and throwing the timing away.
Read handwritingYou want handwritten or cursive text turned into characters.Handwriting recognition: TrOCR or Kraken.These are trained on handwriting specifically, where general OCR struggles.The trap: Pointing printed-text OCR at cursive and wondering why it fails.
Pricing and elasticityYou want to set prices from how demand actually responds, not from a guess.Estimate price elasticity with causal methods, then set prices with an optimization under your margin and volume goals.Elasticity is a causal question, what happens to demand if you move price, and pricing is then an optimization against it. Two different tools, chained.The trap: Asking a language model for a price. It has no read on your demand curve and no way to optimize against your constraints.
Build a knowledge graphYou want to turn documents and records into a connected graph of entities and how they relate.Named entity recognition plus relation and event extraction, feeding a graph.Extract the entities, then the relationships between them, and the connected structure is what you query.The trap: Extracting entities and stopping. Without the relationships you have a list, not a graph.
Which LLM (or AI) should you use?First check whether the job is actually a language job, because most business problems are not. If it genuinely is, the choice between the frontier models is real but secondary.For open-ended language and reasoning, any current frontier model (Claude, GPT-class, or Gemini) is strong; choose on cost, latency, context length, and where your data can live. For a narrow, high-volume task, a small model or a fine-tuned classifier beats all of them.The frontier models are close enough that the bigger lever is matching the model type to the job. Reach for a language model for drafting, reasoning, and summarizing; reach for a specialist for forecasting, classifying, extracting, and optimizing.The trap: Picking a chatbot brand before checking whether the problem is even a language problem. That is how you pay frontier prices to do logistic regression, badly.
The whole field guide (145 approaches, 14 families)
Classical and tabular ML12 approaches

Still the majority of the machine learning that actually runs in production at real companies.

  • Linear, ridge, lasso, elastic netgives you a predicted number

    Draws the straight-line relationship between your inputs and a number, and reads a prediction off it.

    Use when Small to medium data with a roughly straight-line relationship, when you need to point to exactly which factor drove the number.

  • Logistic regressiongives you a category

    Weighs the factors and outputs the probability that something is true or false.

    Use when The honest baseline for any yes-or-no question, and the standard anywhere a decision must be justified one factor at a time.

  • Decision tree (CART)gives you a category

    Learns a flowchart of yes-or-no questions that a person can read and follow by hand.

    Use when You want plain if-this-then-that rules a human can execute, audit, or argue with.

  • Random forestgives you a category

    Asks hundreds of slightly different decision trees and takes their majority vote.

    Use when A strong, low-maintenance default that copes with messy, mixed, and missing data without much tuning.

  • Gradient-boosted trees: XGBoost, LightGBM, CatBoostgives you a predicted number

    Builds decision trees one after another, each fixing the mistakes of the last.

    Use when The default winner on spreadsheet-shaped data. Reach for this first unless you have a specific reason not to.

  • Generalized additive models (GAM, EBM)gives you a predicted number

    Captures curved relationships while still showing a readable graph of how each factor moves the answer.

    Use when You need near-top accuracy and a regulator or clinician must see exactly how each variable affects the result.

  • Naive Bayesgives you a category

    Counts how often each clue goes with each outcome and multiplies the odds together.

    Use when A very fast, very cheap baseline, especially on text and small data.

  • k-nearest neighborsgives you a category

    Predicts by finding the most similar past examples and copying their answer.

    Use when Small datasets, or as a quick lookup over a set of similarity fingerprints.

  • Support vector machine (SVM, SVR)gives you a category

    Draws the widest possible dividing line between two groups.

    Use when High-dimensional data with very few examples, like text or lab spectra.

  • Gaussian processesgives you a predicted number

    Predicts a value and, just as importantly, how sure it is about that value.

    Use when Data is scarce or expensive to collect and you need honest confidence ranges, not just a guess.

  • Bayesian networksgives you found and organized information

    A map of which things cause or depend on which, used to reason about probabilities.

    Use when You have expert knowledge of how factors relate and need to reason under uncertainty.

  • AutoML: AutoGluon, H2O, FLAMLgives you a category

    Automatically tries many models and settings and hands you the best one.

    Use when Day one of a spreadsheet project, to set a strong benchmark before you invest in anything custom.

Unsupervised learning and structure14 approaches

No labels. You are asking the data what shape it already has.

  • K-meansgives you a group

    Sorts records into a set number of groups by how close together they are.

    Use when Fast, simple grouping when you can reasonably guess how many groups to expect.

  • DBSCAN and HDBSCANgives you a group

    Finds groups as dense clumps of points and leaves the stragglers as outliers.

    Use when Groups of odd shapes, an unknown number of groups, and you want outliers flagged too.

  • Gaussian mixture modelgives you a group

    Groups records but lets each one partly belong to several groups at once.

    Use when Overlapping groups where you want a probability of membership, not a hard label.

  • Hierarchical clusteringgives you a group

    Builds a family tree of similarity, from every record alone up to one big group.

    Use when You want the whole tree of relationships, not a single fixed set of groups.

  • PCA and SVDgives you found and organized information

    Squeezes many columns down to a few that still carry most of the information.

    Use when Compress, de-noise, and simplify data before another model uses it.

  • Non-negative matrix factorizationgives you found and organized information

    Breaks data into a handful of additive building blocks you can interpret.

    Use when You want parts that add up to the whole, like topics in text or ingredients in a signal.

  • Independent component analysisgives you found and organized information

    Untangles several signals that got mixed together into one recording.

    Use when Separating mixed sources, such as isolating brain signals from noise.

  • t-SNE and UMAPgives you found and organized information

    Flattens complex data onto a 2-D map so you can see clusters with your eyes.

    Use when Visualization and exploration only. Never feed its output into another model.

  • Latent Dirichlet allocationgives you a group

    The classic way to discover the hidden topics running through a pile of documents.

    Use when A large text collection where you want readable topics and have no GPU.

  • BERTopicgives you a group

    Groups documents by meaning and auto-labels each cluster with its keywords.

    Use when Modern topic discovery. Better than the classic approach in almost every case.

  • Association rules: Apriori, FP-Growthgives you found and organized information

    Finds patterns like people who bought A also bought B.

    Use when Shopping baskets and co-occurrence, when you want explicit rules you can read.

  • Isolation Forest, one-class SVM, LOFgives you a category

    Learns what normal looks like and flags whatever does not fit, with no examples of bad.

    Use when Catching the unusual when you have no labelled examples of fraud or faults, which is most of the time.

  • Autoencoder for anomaly detectiongives you a category

    Learns to rebuild normal data; whatever it rebuilds badly is the anomaly.

    Use when Spotting the odd one out when normal behaviour is complex and high-dimensional.

  • Self-organizing mapgives you a group

    Arranges data on a 2-D grid so similar things sit near each other.

    Use when A niche, older tool for visual clustering in industrial monitoring.

Time series and forecasting13 approaches

A completely separate discipline from language models, and one of the highest-return areas in a real business.

  • ARIMA, SARIMA, SARIMAXgives you a predicted number

    Forecasts by learning a series' own momentum, its seasonal pattern, and outside drivers.

    Use when A handful of series, a stable pattern, and you need statistical rigour and real confidence intervals.

  • Exponential smoothing, Holt-Winters, ETSgives you a predicted number

    Forecasts from a weighted average of recent history, trend, and season.

    Use when A fast, sturdy baseline. Start here before anything more complex.

  • Prophetgives you a predicted number

    A forecaster built for analysts that handles trends, seasons, and holidays out of the box.

    Use when Business forecasts with known holiday and promotion spikes.

  • GARCHgives you a predicted number

    Forecasts how much a number will swing, not where it will land.

    Use when Financial risk and any series where the size of the ups and downs is what matters.

  • Kalman filter and state-space modelsgives you a predicted number

    Continuously updates its best estimate of a hidden value as noisy new readings arrive.

    Use when Live tracking from noisy sensors, like position from GPS.

  • Gradient boosting on lagged featuresgives you a predicted number

    Turns forecasting into a spreadsheet problem by feeding past values in as columns.

    Use when Many series with lots of extra context. The quiet winner of most real forecasting contests.

  • DeepARgives you a predicted number

    Learns one forecasting model across thousands of related series at once, with ranges.

    Use when Hundreds or thousands of related series where you want full probability ranges.

  • N-BEATS and N-HiTSgives you a predicted number

    Deep-learning forecasters that need no manual feature work.

    Use when Strong single-series accuracy with minimal setup.

  • Temporal fusion transformergives you a predicted number

    A forecaster that weighs many inputs and shows which ones drove the prediction.

    Use when Complex forecasts with many drivers where you must explain what moved the number.

  • PatchTST and iTransformergives you a predicted number

    Transformer forecasters tuned for very long histories and many series.

    Use when Long-horizon, many-variable forecasting like grid load or sensor fleets.

  • Foundation models: TimesFM, Chronos, Moiraigives you a predicted number

    Pre-trained forecasters that work on a brand-new series with no training.

    Use when You need a fast forecast on a new series and have no time to train one.

  • Survival analysis: Kaplan-Meier, Cox, DeepSurvgives you a predicted number

    Predicts how long until an event, correctly handling cases that have not happened yet.

    Use when The question is when, not whether: churn timing, equipment failure, contract renewal.

  • Change-point detection: PELT, BOCPDgives you a category

    Spots the moment a trend genuinely shifts, separate from normal noise.

    Use when Deciding whether a metric really changed on a given date, or you are chasing noise.

Neural architectures16 approaches

The primitives. Most of the named products in this guide are one of these wearing a costume.

  • Multilayer perceptrongives you a predicted number

    The basic neural network: layers of simple units that together learn any pattern.

    Use when The final decision layer sitting on top of a bigger model. Rarely the whole answer by itself.

  • Convolutional networks: ResNet, EfficientNet, ConvNeXtgives you a category

    A neural network that scans an image in small patches to recognize shapes and textures.

    Use when Image tasks that must be small, fast, or run on a device. Still the go-to for everyday vision.

  • RNN, LSTM, GRUgives you a predicted number

    A network that reads a sequence one step at a time, carrying a memory of what came before.

    Use when Sequences on tiny hardware or under strict live-streaming limits.

  • Transformergives you found and organized information

    The architecture behind modern AI: it weighs every part of the input against every other part.

    Use when An architecture, not a product. Vision models, speech models, forecasters, and AlphaFold are all transformers, and none of them are language models.

  • Vision transformergives you a category

    A transformer that treats an image as a grid of patches instead of words.

    Use when Large image datasets where you want strong transfer learning. Needs more data than a CNN.

  • Graph neural networks: GCN, GraphSAGE, GATgives you a category

    A network that learns from how things are connected, passing information along the links.

    Use when Your data is a network of relationships and the connections carry the signal.

  • Autoencoder and VAEgives you found and organized information

    Squeezes data down to a compact code and rebuilds it, learning the essence in between.

    Use when De-noising, anomaly detection, and compression. The compact code is the useful part.

  • Generative adversarial network (GAN)gives you newly generated content

    Two networks compete: one forges, one detects, until the forgeries look real.

    Use when Fast generation, realistic fake tabular data, and image upscaling. Largely replaced by diffusion for images.

  • Diffusion modelsgives you newly generated content

    Starts from pure noise and repeatedly cleans it up into a realistic image or signal.

    Use when The highest-quality generation across images, video, audio, and molecules.

  • Flow matching and rectified flowgives you newly generated content

    A newer generator that learns a direct path from noise to a finished result.

    Use when Modern image generation that reaches quality in far fewer steps than diffusion.

  • Normalizing flowsgives you a predicted number

    A generator that can also state the exact likelihood of any given sample.

    Use when You need an exact probability of a sample, not just the ability to produce one.

  • State-space models: Mamba, S4gives you found and organized information

    Processes very long sequences efficiently, in one pass, without attention's cost.

    Use when Extremely long inputs, live streaming, and tight efficiency budgets.

  • Mixture of expertsgives you found and organized information

    Routes each piece of input to a few specialist sub-networks instead of the whole model.

    Use when Growing a model's capacity without growing the cost of each prediction. Most frontier LLMs work this way.

  • NeRF and 3D Gaussian splattinggives you newly generated content

    Turns a handful of ordinary photos into an explorable 3-D scene.

    Use when Digital twins, as-built capture, property walkthroughs, and visual effects.

  • Physics-informed neural networksgives you a predicted number

    A network that is forced to obey known physical laws while it learns.

    Use when Sparse data plus known physics: fluid flow, structures, subsurface modelling.

  • Siamese and contrastive networksgives you a category

    Learns to tell whether two things are the same or different.

    Use when Verification and similarity when you have very few examples of each item, like matching signatures.

Computer vision19 approaches

Organized by task, not by model. This is the part of AI that has nothing to do with language at all.

  • Image classification: EfficientNet, ConvNeXt, ViTgives you a category

    Gives a whole image one label.

    Use when The entire image gets a single verdict, like pass or fail.

  • Object detection: YOLO, RT-DETR, Grounding DINOgives you found and organized information

    Draws a box around each object and says what it is.

    Use when You need to know what is in the frame, where, and how many.

  • Segmentation: U-Net, Mask R-CNN, SAM 2gives you found and organized information

    Traces the exact outline of each object, down to the pixel.

    Use when You need the precise shape or area, not just that something is present.

  • Pose and keypoint estimation: MediaPipe, RTMPosegives you found and organized information

    Tracks the position of body, hand, and joint points.

    Use when Movement and posture are the signal: ergonomics, physical therapy, sports.

  • Face detection and recognition: RetinaFace, ArcFacegives you a category

    Finds faces and matches them to identities.

    Use when Access control and identity. Gate this hard on consent and law.

  • OCR: Tesseract, PaddleOCR, Surya, TrOCRgives you found and organized information

    Reads printed text out of an image into characters.

    Use when Any time text is trapped inside a scan or photo.

  • Document AI: LayoutLMv3, Donut, Nougatgives you found and organized information

    Reads a document and understands which text is which field, plus tables.

    Use when You need clean structured data out of forms and invoices, not a wall of text.

  • Handwriting recognition: TrOCR, Krakengives you found and organized information

    Reads handwritten and cursive text.

    Use when Historical records, clinical notes, and handwritten field forms.

  • Object tracking: ByteTrack, DeepSORTgives you found and organized information

    Follows the same object from frame to frame across a video.

    Use when Queue times, dwell time, throughput, and sports analytics.

  • Optical flow: RAFTgives you found and organized information

    Measures how each pixel moves between video frames.

    Use when When motion itself is what you are measuring.

  • Depth estimation: Depth Anything, MiDaSgives you found and organized information

    Estimates how far away things are from a single flat photo.

    Use when Robotics, AR, and taking measurements from an ordinary picture.

  • Visual anomaly detection: PatchCore, PaDiMgives you a category

    Learns what a good part looks like and flags anything different.

    Use when Factory quality control, where you have good samples but not every possible defect.

  • Super-resolution and restoration: Real-ESRGAN, SwinIRgives you newly generated content

    Sharpens, upscales, and repairs degraded images.

    Use when The source is low-quality and the detail matters.

  • Zero-shot vision: CLIP, SigLIPgives you a category

    Finds or labels images from a plain text description, with no training.

    Use when A huge image library with no labels: find every photo with a forklift.

  • Image generation: Stable Diffusion, Flux, Imagen, Ideogramgives you newly generated content

    Creates a new image from a text description.

    Use when You need a picture that does not exist. Ideogram when the text inside the image must be legible.

  • Conditioning and control: ControlNet, LoRA, IP-Adaptergives you newly generated content

    Steers image generation to hold a pose, layout, or brand style.

    Use when Generation alone is not enough; you need it on-brand and on-composition.

  • Video generation: Sora, Veo, Runway, Klinggives you newly generated content

    Creates moving video from a text prompt or a still image.

    Use when You need footage you cannot practically shoot.

  • Video understanding: VideoMAE, TimeSformergives you a category

    Recognizes what is happening across a stretch of video.

    Use when The event unfolds over time; a single frame will not tell you.

  • Multimodal LLMgives you found and organized information

    A language model that can also look at an image and reason about it.

    Use when The question about an image is open-ended and needs judgement, not a fixed label.

Audio and speech10 approaches

An entire modality with nothing to do with text generation.

  • Speech recognition: Whisper, Parakeet, Conformergives you found and organized information

    Turns spoken words into written text.

    Use when Any spoken input: calls, meetings, voice notes.

  • Speaker diarization: pyannotegives you found and organized information

    Works out who spoke and when.

    Use when More than one person is talking and you need it attributed.

  • Speaker verification: ECAPA-TDNNgives you a category

    Confirms whether a voice belongs to a specific person.

    Use when Verifying identity from a voice.

  • Text to speech: ElevenLabs, Kokoro, Pipergives you newly generated content

    Reads written text aloud in a natural voice.

    Use when The output needs to be heard, not read: phone systems, accessibility, narration.

  • Voice conversion: RVC, So-VITSgives you newly generated content

    Changes whose voice a recording sounds like.

    Use when Dubbing and localization. Consent and law gate this, not the technology.

  • Audio event classification: YAMNet, PANNs, ASTgives you a category

    Recognizes non-speech sounds.

    Use when The signal is a sound, not a word: a failing bearing, breaking glass, an alarm.

  • Music generation: MusicGen, Stable Audiogives you newly generated content

    Generates music and sound effects.

    Use when You need original audio or a background track.

  • Source separation: Demucs, Spleetergives you newly generated content

    Splits a mixed recording into separate tracks.

    Use when You need one element, like a single voice, out of a noisy mix.

  • Keyword spottinggives you a category

    Listens continuously for a specific wake word using almost no power.

    Use when Always-on, battery-powered listening for a wake word.

  • Audio embeddings: CLAP, VGGishgives you found and organized information

    Turns sounds into searchable fingerprints.

    Use when Searching or grouping audio by how it sounds, like every clip of a failing bearing.

Language models (LLM)6 approaches

One family out of fourteen. The front door of the building, not the building.

  • Frontier LLM: Claude, GPT-class, Geminigives you newly generated content

    The large, general models that write and reason over language.

    Use when Hard reasoning, agent workflows, and genuinely open-ended language.

  • Reasoning models (extended thinking)gives you newly generated content

    A language model that deliberately thinks longer before it answers.

    Use when Multi-step maths, planning, proofs, and difficult code.

  • Small language models: Haiku-class, Phi, Gemmagives you newly generated content

    Compact, cheaper language models that still handle everyday language tasks.

    Use when High volume, low latency, on-device, or cost-sensitive work.

  • Code models: Claude Code, Qwen-Codergives you newly generated content

    Language models specialised in writing, fixing, and reviewing code.

    Use when The output is source code.

  • Domain fine-tune (LoRA, SFT)gives you newly generated content

    A base model further trained on your own examples to specialise it.

    Use when A narrow, repeated task where you already have labelled examples.

  • Encoder-decoder: T5, BART, Pegasusgives you newly generated content

    Language models built to transform one text into another, predictably.

    Use when Summarize, translate, or restructure text. A transform, not a conversation.

Language work without generation10 approaches

Cheaper, faster, and usually more accurate than a language model. The boring, profitable half of NLP.

  • Text embeddings: BGE, E5, Voyage, Coheregives you found and organized information

    Turns text into number-fingerprints so a computer can measure meaning and similarity.

    Use when The engine under search, retrieval, clustering, de-duplication, and find-similar.

  • Sparse retrieval: BM25, SPLADEgives you a ranking or recommendation

    Classic keyword search that ranks documents by exact word matches.

    Use when Part numbers, IDs, and names, exactly where meaning-based search quietly fails. Pair it with embeddings.

  • Rerankers: Cohere Rerank, BGE-rerankergives you a ranking or recommendation

    Takes a shortlist of results and re-orders it for a much sharper top few.

    Use when The single highest-value upgrade to any search or retrieval system.

  • Text classification: DeBERTa, DistilBERT, fastTextgives you a category

    Reads text and drops it into predefined categories, at scale and cheaply.

    Use when High-volume labelling: tickets, intent, sentiment, moderation. These understand text without generating it, at a fraction of an LLM's cost.

  • Named entity recognition: spaCy, GLiNERgives you found and organized information

    Pulls the names, dates, and amounts out of free text.

    Use when You need specific fields extracted, not a summary.

  • Relation and event extractiongives you found and organized information

    Finds how the entities in text relate, and links them up.

    Use when Turning documents into a connected knowledge graph.

  • Machine translation: NLLB, MarianMT, SeamlessM4Tgives you newly generated content

    Translates text between languages.

    Use when Bulk translation, where paying an LLM per word would be ruinous.

  • Extractive question answeringgives you found and organized information

    Answers a question by quoting the exact sentence from a document, inventing nothing.

    Use when Compliance and legal answers where a made-up answer is unacceptable.

  • Safety and moderation classifiers: Llama Guardgives you a category

    Screens text and blocks unsafe or policy-violating content.

    Use when Any system open to the public or to untrusted input.

  • Grammar and style: LanguageToolgives you a category

    Checks writing against fixed grammar and style rules.

    Use when You want consistency and correctness, not creativity.

Recommender systems7 approaches

Trained on your behavioral log, which a language model has never seen.

  • Matrix factorization: ALS, SVD++gives you a ranking or recommendation

    Learns hidden taste patterns from who interacted with what.

    Use when You have an activity log and few details about the items themselves.

  • Content-based filteringgives you a ranking or recommendation

    Recommends items similar to ones a person already liked, by their features.

    Use when New items with rich descriptions and little interaction history yet.

  • Two-tower retrievalgives you a ranking or recommendation

    Instantly narrows millions of items to a relevant shortlist for each person.

    Use when Stage one of every modern recommender: a thousand candidates from millions, in milliseconds.

  • Ranking models: Wide and Deep, DeepFM, DLRMgives you a ranking or recommendation

    Scores each shortlisted item by how likely the person is to click or buy.

    Use when Stage two: putting the shortlist in the best order.

  • Sequential recommenders: SASRec, BERT4Recgives you a ranking or recommendation

    Predicts the next item from the order of what someone did before.

    Use when Session-based what-comes-next, like the next episode or track.

  • Graph recommenders: LightGCN, PinSagegives you a ranking or recommendation

    Recommends by spreading signals across the web of users and items.

    Use when Sparse activity with strong word-of-mouth or network effects.

  • Learning to rank: LambdaMARTgives you a ranking or recommendation

    Trains directly on the quality of the final ordering.

    Use when When getting the order of results right is the actual job, like site search.

Reinforcement learning and decisioning7 approaches

Reinforcement learning is powerful and hard. Bandits are cheap, safe, and badly underused. Start with bandits.

  • Multi-armed bandit: Thompson sampling, UCBgives you a decision or plan

    Tries several options and steadily shifts toward the winner while still testing the rest.

    Use when Live experiments where you want to stop wasting traffic on the losing option: headlines, layouts, offers.

  • Contextual banditgives you a decision or plan

    Like a bandit, but it looks at who each person is, so the best option can differ per person.

    Use when Personalized choices with instant feedback, like which offer to show this particular visitor.

  • Q-learning and DQNgives you a decision or plan

    Learns by trial and error which action pays off in each situation, the way you learn a game by playing it.

    Use when Step-by-step decision problems where you can practise safely in a simulator first.

  • Policy gradient: PPO, SAC, TD3gives you a decision or plan

    Learns a strategy for smooth, continuous control like steering or throttle.

    Use when Robotics and control problems trained against a simulator.

  • Model-based RL and MuZerogives you a decision or plan

    Builds its own model of how the world reacts, then plans ahead inside it.

    Use when When real trial and error is too slow or costly, so learning efficiency matters.

  • Offline RL and decision transformersgives you a decision or plan

    Learns the best sequence of actions from a fixed history, never experimenting live.

    Use when When you cannot safely experiment on real customers or patients, so you learn from records.

  • RLHF, DPO, GRPOgives you a decision or plan

    Tunes a generative model using human preferences about which answers are better.

    Use when Post-training a language model to be helpful, safe, and on brand.

Optimization and operations research7 approaches

Not AI in the marketing sense, and frequently the actual answer.

  • Mixed-integer programming: OR-Tools, Gurobi, CPLEXgives you a decision or plan

    Finds the provably best plan that never breaks a hard rule.

    Use when Scheduling, routing, and budgets with constraints that cannot be violated.

  • Constraint programming: CP-SATgives you a decision or plan

    Fits many rigid requirements together into a workable schedule.

    Use when Rostering, timetabling, and assignment puzzles.

  • Vehicle routing solversgives you a decision or plan

    Plans the most efficient routes under delivery windows and vehicle limits.

    Use when Fleets, deliveries, and field-service dispatch.

  • Genetic algorithms and simulated annealinggives you a decision or plan

    Searches a huge space of options by mutating and improving candidate solutions.

    Use when The goal is a black box or has no clean formula: layout, design, tuning.

  • Bayesian optimizationgives you a decision or plan

    Chooses the next experiment to run so you learn the most from the fewest trials.

    Use when Each trial is expensive: lab work, physical tests, or model tuning.

  • Monte Carlo simulationgives you a predicted number

    Rolls the dice thousands of times to show the full range of outcomes.

    Use when You need a risk range and odds, not a single point estimate.

  • Discrete-event and agent-based simulationgives you a predicted number

    Builds a working model of a system so you can test changes before they are real.

    Use when Try a policy on a simulated warehouse, queue, or market before deploying it.

Causal inference8 approaches

Predictive models tell you what will happen. These tell you what will happen if you act. That difference is worth a great deal of money.

  • Randomized experiment (A/B test)gives you a predicted number

    Splits people at random to measure the true effect of a change.

    Use when Whenever you can genuinely randomize. This is the gold standard.

  • Uplift and CATE: causal forest, T-learner, X-learnergives you a decision or plan

    Estimates how much a treatment changes each individual, not just the average.

    Use when Targeting the people an offer will actually sway, rather than those who would act anyway.

  • Propensity score matchinggives you a predicted number

    Builds a fair comparison group from data you already have.

    Use when You could not run an experiment but still need a credible cause-and-effect claim.

  • Difference in differencesgives you a predicted number

    Compares the before-and-after of a treated group against an untreated one.

    Use when A change rolled out to some regions or teams and not others.

  • Synthetic controlgives you a predicted number

    Builds a stand-in for what would have happened from a blend of similar cases.

    Use when One treated unit, like a single store or market, and many untreated ones.

  • Instrumental variablesgives you a predicted number

    Isolates a true effect when hidden factors are muddying the data.

    Use when Hidden confounding is distorting a plain comparison.

  • Marketing mix modeling: Robyn, Meridiangives you a predicted number

    Estimates how much each channel actually contributes to sales.

    Use when Privacy-safe, top-down decisions about where to put the marketing budget.

  • Structural causal models: DoWhy, EconMLgives you found and organized information

    Writes your cause-and-effect assumptions as an explicit, testable diagram.

    Use when High-stakes does-X-cause-Y claims you will have to defend.

Scientific and domain models9 approaches

Where AI stopped being a chatbot and started being an instrument.

  • Protein structure: AlphaFold 3, ESMFold, Boltzgives you newly generated content

    Predicts a protein's 3-D shape from its sequence.

    Use when Structural biology and drug discovery.

  • Molecular models: property GNNs, generative design, dockinggives you newly generated content

    Predicts how molecules behave and designs new ones.

    Use when Chemistry and pharmaceutical screening.

  • Weather: GraphCast, Aurora, GenCastgives you a predicted number

    Forecasts weather as well as supercomputer models, far faster and cheaper.

    Use when Energy, agriculture, logistics, and insurance planning.

  • Materials: GNoME, M3GNetgives you newly generated content

    Discovers new stable materials before anyone makes them.

    Use when Materials science and battery research.

  • Geospatial: Prithvi, satellite segmentationgives you found and organized information

    Reads satellite and aerial imagery to map what is on the ground.

    Use when Land use, deforestation, crop yield, and damage assessment.

  • Medical imaging: nnU-Net, MedSAMgives you found and organized information

    Finds and outlines structures in CT, MRI, and X-ray scans.

    Use when Clinical imaging and screening support.

  • Genomics: Enformer, DNABERT, Evogives you a predicted number

    Reads DNA to predict what genes do and how variants matter.

    Use when Genomics research and interpreting variants.

  • Robotics: RT-2, OpenVLA (vision-language-action)gives you a decision or plan

    Turns what a robot sees plus an instruction directly into movement.

    Use when Machines that must perceive and act in the physical world.

  • Autonomous driving: BEVFormer, occupancy networksgives you a decision or plan

    Turns camera and sensor feeds into a driving picture and a plan.

    Use when Self-driving and driver-assistance systems.

The supporting cast7 approaches

Not models, but you cannot ship the models without them.

  • Vector databases: pgvector, Qdrant, Pineconegives you a ranking or recommendation

    Stores those text fingerprints and finds the closest matches instantly.

    Use when The moment you use embeddings, they need somewhere to live and be searched.

  • Feature stores: Feast, Tectongives you found and organized information

    Serves the exact same inputs to training and to the live model.

    Use when Prevents the most common silent production bug: training and serving seeing different data.

  • Quantization and distillation: GPTQ, AWQ, GGUFgives you found and organized information

    Shrinks a model to run cheaper, faster, or on smaller hardware.

    Use when The model works but you cannot afford to run it as it stands.

  • Synthetic data: CTGAN, SDVgives you newly generated content

    Generates realistic fake data that carries no private information.

    Use when Real data is restricted, or a rare case is badly under-represented.

  • Explainability: SHAP, LIMEgives you found and organized information

    Shows which inputs drove a given prediction.

    Use when Required for regulated decisions, and your best debugging tool everywhere else.

  • Drift monitoring: Evidently, NannyMLgives you a category

    Watches for when live data drifts away from what the model was trained on.

    Use when Models silently decay; you want to catch it before the business does.

  • Guardrail modelsgives you a category

    Screen and block bad inputs and outputs at the edge of the system.

    Use when Any system exposed to untrusted input.

A note on what is in here, and what is not

The guide names techniques and libraries rather than model releases, which is deliberate: gradient-boosted trees, DeBERTa, fastText, HDBSCAN, Isolation Forest, BM25, Whisper, PatchCore, PaDiM, constraint solvers, uplift and CATE modeling. None of those go stale the way a version number does.

The only frontier-model names anywhere in it are search keywords, so that someone typing the question the way they would say it out loud still finds the guide. They are there to catch a query, not to make a claim about which product is best. We do not sell software and we take nothing from anyone whose product we might recommend.

The field guide was assembled in July 2026. If you think something is missing from it, tell us and we will look.

Work with us

Tell us what your week looks like. We will tell you honestly whether an hour with us is worth your time.

Book a first session