When you bought your last press brake, you started with the parts you run: tonnage, bed length, the thickest plate you bend on a bad day. Then you looked at who could service it, what it cost to run, and whether your operators could learn it. The logo on the side came last.

Buying AI should work the same way, but most people start with the logo. They hear a new model name every few weeks, see a headline about a record test score, and assume they are already behind. They are not, and this guide covers who makes the models, what each is known for, and what the test numbers can and cannot tell you about your work.

73%drop in Anthropic's top-tier list price, Opus 4.1 (2025) to Opus 5.5Anthropic pricing
47%fall per quarter in the cost of a given AI performance level since 2023Epoch AI
41.7%of long desktop workflows finished by the best reported model (20.6% in July)Anthropic, OSWorld 2.0

What an AI model is

A model is the AI software a company trains on huge amounts of text, code, and images so it can read and write. A handful of companies, usually called labs, build the best ones. Everything else you hear about (the chat apps, the AI inside your ERP, the quoting add-on at the trade show) runs on top of a model from one of these labs.

A model knows nothing about your company until you show it. It does not learn from your corrections on its own either. The examples, instructions, and approval rules that make it useful to you live in files it reads every time it works, and those files are yours.

The labs and their current models

Models come in two kinds, and the difference matters to anyone with an NDA or ITAR work.

You rent closed models. The lab keeps the model on its own servers, you send your text in and get an answer back, and you pay by the amount of text. You never hold the model itself.

You can own a copy of an open-weight model. The lab publishes the model's "weights" (the trained settings that make it work), and anyone can download them and run them on their own hardware, so nothing leaves the building. The very best models are still closed, though, and running your own takes equipment and someone who knows how.

Closed models you rent from the lab

Anthropic makes Claude. Its top generally available model is Claude Fable 5.1, released September 1. A sibling, Mythos 5.1, is the same model with different safeguards and is limited to vetted US organizations doing cybersecurity and life sciences work. Three weeks later Anthropic released Claude Opus 5.5, which it says performs at Fable 5.1 level on most work for less money. Anthropic is known for coding, for agents (AI that carries out multi-step tasks instead of only answering), and for office deliverables like spreadsheets, slides, and documents.

OpenAI makes GPT and ChatGPT. Its top model is GPT-6 Astra, released September 3, which OpenAI calls its most capable broadly deployed model. On September 22 it added two cheaper siblings, GPT-6 Sol and GPT-6 Luna. OpenAI is known for ChatGPT's huge user base, reasoning and math, and its Codex coding agent.

Google makes Gemini. Its top model is still Gemini 3.1 Pro, from February. Google says it hopes to release Gemini 4 well before the end of 2026 but has given no date. Meanwhile it keeps shipping cheaper, faster "Flash" models, most recently Gemini 3.8 Flash on September 2. Google is known for handling video, audio, and images, and for putting Gemini inside Workspace and Android, which you may already pay for.

Meta plays both sides. Muse Spark, released in April, replaced Llama as its flagship and is closed. xAI, the company behind Grok, is also closed. We could not confirm its latest model or price from xAI itself, so we leave it there.

Open-weight models you can run on your own hardware

Read this list if your drawings cannot go to someone else's server.

  • Meta Muse Glimmer, released August 10, is built to run locally under a permissive license (Apache 2.0, which lets you use it commercially).
  • DeepSeek, from China, released V4 Pro and V4 Flash in April under the permissive MIT license.
  • Alibaba's Qwen, also from China, released Qwen3.8-27B in August. It runs on laptop-class hardware under Apache 2.0.
  • Mistral, from Europe, released Mistral Medium 3.5 as open weights in April. The same changelog shows Mistral's document-reading model, OCR 4.1, reached general availability on August 31, and reading MTRs, POs, and drawings is that kind of work.

Where a model comes from may matter to your customers, especially defense customers, so ask before you pick. A license is also a contract. Qwen's largest open model, for example, requires a separate agreement for providers with more than $50 million in revenue, so read the license the way you would read a customer's terms and conditions.

What AI test scores (benchmarks) measure

A benchmark is a standard test that every model takes so you can compare them. It works like a weld certification: the same coupon and the same bend test, scored the same way, so you can compare two welders who have never met.

Models keep acing the old tests. In 2024 and 2025, the popular tests covered things like PhD-level science questions and fixing small bugs in code, and the top models now score so high on those that the tests no longer tell them apart. OpenAI stopped reporting one of the best-known coding tests in February, citing flawed test questions and answers that had leaked into training data.

The tests that still separate the leaders measure long, multi-step work, which is closer to what you would hire AI to do:

  • Terminal-Bench checks whether an AI can finish multi-step jobs on a computer, like setting up software or processing data. Anthropic reports Opus 5.5 at 66.4%, up from 52.3% for Opus 5.
  • OSWorld 2.0 checks whether an AI can finish long office workflows (about 1.6 human hours each, on median) by clicking and typing in real desktop apps. When it launched in July, the best score was 20.6%. By September, Anthropic reported 41.7% for Fable 5.1. That is real progress, but the model still failed more of these jobs than it finished.
  • GDPval has experts grade real work products (documents, spreadsheets, slides) from 44 occupations, including manufacturing.

Two scorecards try to sum it up. Artificial Analysis runs the same set of tests on every model, which makes its index one of the few scores you can compare across labs.

Artificial Analysis Intelligence Index, leading models from three labs index score, each model at its maximum setting
Claude Opus 5.558
Claude Fable 5.153
GPT-6 Astra53
Meta Muse Spark 1.348

Source: Artificial Analysis leaderboard, read 2026-09-25.

Meta's best model ranks 13th on the full list, and it still sits only 10 points behind the leader. The top scorer, Opus 5.5, also costs less than half as much to use: $4 in and $20 out per million tokens (prices are explained below), against $10 and $50 for Fable 5.1 and GPT-6 Astra. Our view: for a manufacturer, a few index points matter less than price, privacy, and what your other software already uses, because none of the index tests looks like your paperwork.

Outside that index, labs choose which tests to publish and often run them their own way, so scores from different vendors usually cannot be compared. Some widely repeated GPT-6 Astra scores appear only in secondary coverage with different setups, so we leave them out.

What a benchmark does not tell you

A weld cert tells you a welder can make a good joint on a test plate. It says nothing about whether they show up on time, read your drawings right, or know your customer wants the grind marks facing in.

Benchmarks have the same limit. No benchmark has seen your part numbers, your customers' cert requirements, your ERP, or the way your estimator handles a revision on a drawing that came in as a scanned PDF. The score tells you a model can do hard work in general, and only a test on your own documents tells you whether it can do yours.

That test is called a test set: 50 real past cases of one job with the approved answer attached, set aside before anyone builds anything and used to judge every model and every version. For a cert packet, you score it field by field (heat number, grade, quantity) and mark which fields count as critical errors, where one wrong value ships a bad part.

The benchmark is the cert. Your first article is still your job.

What AI models cost, and where prices are going

Models are priced per token, a small chunk of text, roughly a piece of a word. Prices are quoted per million tokens, with a lower rate for what you send in and a higher rate for what comes back.

Prices are falling fast. Anthropic's top-tier list price went from $15 in and $75 out per million tokens for Claude Opus 4.1 in 2025 to $4 and $20 for Opus 5.5, about 73% lower. OpenAI cut GPT-6 Sol to $2 and $10, half the old price. Its smallest model, Luna, costs 10 cents per million tokens in.

The research group Epoch AI estimates that the cost of a given level of AI performance has fallen about 47% per quarter since 2023, roughly 13 times cheaper every year. Its example: a science test question that cost about 30 cents to answer at a certain level in early 2025 cost about four hundredths of a cent later with a smaller model.

The very best models still cost real money. Running the full Artificial Analysis test suite costs $5.98 on Opus 5.5 and $7.63 on Fable 5.1. An agent working through a long job also uses far more tokens than a quick chat, so your total bill can go up even as the price per token drops, and Anthropic's newer models count about 30% more tokens for the same text.

So budget by the job. Ask any vendor what one finished job costs in model fees (one quote, one cert packet), not the price per million tokens. Our view is that setup and your reviewers' time will cost more than the model bill, because a million output tokens on Opus 5.5 lists at $20, less than an hour of a quality lead's time.

How to avoid getting locked into the wrong model

The reasonable worry is that you will pick the wrong model, get locked in, and watch the whole field change the next month.

The field does change that fast. This month alone, Anthropic released two new models, OpenAI released three, and Google has another major release on the way. In the 52 weeks to September 25, 2026, the eight biggest labs made 52 notable model releases, one a week on average.

Notable model releases by lab, 2025-09-25 to 2026-09-25 new main-line models made generally available
Anthropic12
OpenAI8
Google8
Alibaba Qwen7
xAI6
Meta4
DeepSeek4
Mistral3

Source: Second Shift count from lab announcements, changelogs, and pricing pages. Six of the 52 (three xAI, two Qwen, one Meta) rest on secondary sources only. Without them the total is 46.

At that pace, a plan that depends on picking the permanent winner will fail.

Build the plan around what you own instead: your data, workflows, and approval rules. That means the parts list cleaned up and loaded, the shared inbox mapped by request type, and the rule that a quality lead signs every cert before it goes out. None of that belongs to a lab. When a better or cheaper model comes out, you run your test set on it, and if it clears the same bar, you switch.

You get locked in when a vendor ties your workflow to one model. Ask any vendor three questions:

  1. Can we switch the underlying model without rebuilding the workflow?
  2. Do we own our data, our instructions to the model (prompts), our examples, and our records, and can we export them?
  3. Can the sensitive work run on an open-weight model on our own hardware if a customer requires it?

If the answers are yes, the model brand is a line item you revisit once a year.

Choose a model by fit

Fit matters more than brand. A shop drowning in scanned MTRs cares about document reading, and a shop with ITAR work cares about open weights. A shop that lives in Microsoft or Google tools may get good-enough AI inside what it already pays for.

The test that matters most is your own test set, run through a model and checked by the person who does that work today.

What to do this week

None of these require buying anything.

  1. List the documents that can never leave your network, such as ITAR drawings and customer IP under NDA. That list decides whether you need open-weight models for some work.
  2. Ask your ERP, quoting, and office software vendors which AI features you already have and which model runs underneath.
  3. Start a test set for one repetitive job (cert requests, POs, RFQ emails): pull 50 real past cases with the approved answer attached, mark the fields where a wrong value would cause harm, and put the folder where nobody uses it to set anything up. It will outlast any model.
  4. Ask the three lock-in questions of any vendor who pitches you this quarter, and write down the answers.
  5. Skip next month's model launch and put a note on the calendar to review the lineup once a quarter. Spend the time on your data and workflows instead.

Questions people ask

What is the difference between closed and open-weight AI models?

A closed model stays on the lab's servers, and you rent it by paying for the amount of text you send and receive. An open-weight model has its trained settings published, so you can download it and run it on your own hardware without data leaving the building. As of September 2026, the very best models are still closed.

How much do AI models cost to use?

AI models are priced per million tokens, with a lower rate for text sent in and a higher rate for text that comes back. As of September 2026, Anthropic's Claude Opus 5.5 lists at $4 in and $20 out per million tokens, and OpenAI's GPT-6 Sol at $2 and $10. For most manufacturers, setup and staff time will cost more than the model itself.

How do I avoid AI vendor lock-in?

Ask any AI vendor three questions: can you switch the underlying model without rebuilding the workflow, do you own your data, prompts, and records and can you export them, and can sensitive work run on an open-weight model on your own hardware. If the answers are yes, the model brand becomes a line item you revisit once a year.

Terms in this piece

Model
the AI software a lab trains on huge amounts of text, code, and images so it can read and write, such as Claude, GPT, or Gemini. It knows nothing about your company until you show it.
Lab
a company that builds and trains its own models, such as Anthropic, OpenAI, or Google.
Closed model
a model you rent. It stays on the lab's servers and you pay for what you use.
Open-weight model
a model whose trained settings are published so you can download it and run it on your own hardware.
Benchmark
a standard test every model takes so their results can be compared, like a certification test.
Token
a small chunk of text, roughly a piece of a word. Model prices are quoted per million tokens.
Agent
a model put to work, taking steps such as reading an email, looking up an order, and drafting a document.
Test set
50 real past cases with known answers, sealed before building starts and used to judge every model and every version.