New Report: Leading AI Models Score at Most 52% on Industrial Parts Questions, Reshape Automation Finds
New Report: Leading AI Models Score at Most 52% on Industrial Parts Questions, Reshape Automation Finds
Reshape Automation’s new Industrial AI Accuracy Index 2026 tested three leading AI models, GPT-6 Astra, Claude and Gemini, on 100 questions drawn from what customers ask distributors and OEMs every day, covering parts from 14 industrial manufacturers. From memory alone, the models got 12% to 14% right. With web search, the best scored 52%. About 1 in 3 wrong answers gave a specific part number or figure with no caveat.
One of the 100 industrial parts questions in Reshape Automation’s report: a customer asks a distributor whether a Siemens module works with their soft starter. A general-purpose AI model says yes, missing that the two belong to different product series. ReshapeX answers from the manufacturer’s own documentation, says no and names the right series before a wrong part gets quoted or ordered.
SAN FRANCISCO--(BUSINESS WIRE)--A customer asks their distributor whether a Siemens communication module, part number 3RW5950-0CH00, works with a 3RW52 soft starter, and the distributor turns to AI to find out. Claude, with web search turned on, said yes all three times it was asked. The verified answer is no. The module belongs with the 3RW55 series, and an order placed on that answer ships a part that doesn’t fit the equipment it was bought for.
That question is one of 100 in the Industrial AI Accuracy Index 2026, published today by Reshape Automation, covering 14 manufacturer catalogs including Siemens, Festo, Rittal and ATI Industrial Automation.
What the models got wrong
- Most of it. With web search, GPT-6 Astra scored 52% and Claude 40.5%. Without it, all three models landed between 12% and 14%. Even the best setup got 42 of the 100 questions wrong, with partial answers earning half credit.
- Wrong with confidence. About 1 in 3 wrong answers gave a specific code or figure with no caveat. Asked when a Festo actuator was discontinued, one model gave a different part number and date on each of three tries, and none was correct.
- Different answers to the same question. Each question went to five model setups, three times each, and in 66 of those 500 pairings the runs didn’t agree. Asked for a Rittal KX terminal box in 304 stainless, 200 by 200 by 80, Gemini gave a specific SKU, then said no such box exists in that range, then gave a different SKU.
- Worst on configured parts. Where a part number is built from the manufacturer’s ordering rules instead of printed in a catalog, as with ATI’s robot tool changers, GPT-6 Astra with web search scored 22.7% and Claude scored 0%. On ordinary part lookups, web search had raised scores by 51 points for GPT-6 Astra and 35 for Claude.
Why the models get parts questions wrong
“The AI products on the market today are mostly a general-purpose model with search or a document store bolted on, and they’re very good at a lot of work,” said Juan Aparicio, CEO and co-founder of Reshape Automation. “Parts questions are different. The answer is usually a relationship between three or four facts that live in different places, like a catalog, a datasheet or a cross-reference table, and some of it was never written down at all. If the model has to rebuild that relationship every time someone asks, it’ll sometimes get it right and sometimes not, and it’ll sound equally sure both times.”
“Distributors are where these questions land,” Aparicio said. “A customer asks whether a part fits, someone on the inside sales team has a few minutes to answer, and one wrong digit ships the wrong part. A tool that’s right half the time doesn’t save that person any work, because they still have to check every answer.”
What ReshapeX does differently
ReshapeX is Reshape Automation’s family of AI agents for the technical and commercial work behind industrial sales and support: customer service, application engineering, inside sales and field maintenance. The agents select and size products, build valid part numbers from specs, find replacements and accessories, and cross-reference and price a bill of materials hundreds of lines long. They quote RFQs, answer “where’s my order?”, walk a technician through a fix, and show a sales rep what an account bought, stopped buying and could buy next.
They work inside the ERP and CRM and wherever customers and staff already are: website chat, Outlook, Microsoft Teams, WhatsApp and voice. Reshape customers report 60% to 90% time savings in technical support, customer service and inside sales, and an 18.6% increase in average order value.
The agents leverage leading AI models, such as GPT and Claude, and Reshape builds the harness around the model: the knowledge it answers from, the tools it can call, and the checks an answer passes before a customer sees it.
That knowledge is assembled once, before anyone asks. For each manufacturer, Reshape pulls the catalog, datasheets, cross-reference data and the rules the brand’s engineers know into one knowledge graph that stores the relationships: what fits what, how a configured part number is built, what replaced a discontinued part. Every fact carries its source, and engineers review the graph before any agent uses it. The agents still search the web or check live stock when that’s the right source. Ask the same question twice and you get the same answer.
item, the German maker of aluminum building-kit systems, runs a ReshapeX agent on its U.S. and Mexico websites in English and Spanish. The agent answers from item’s own product data across more than 4,500 mutually compatible components, and when a question goes past what it can verify, it hands the customer to an item Expert with the whole conversation attached.
“Most AI on a website waits for you to ask and forgets you the moment you leave. The item agent is part of item’s journey, not a box bolted onto it,” said Germain Dufossé, CEO of item America.
How the test was run
Reshape’s engineers wrote up the 100 questions from what customers ask the distributors and OEMs Reshape works with, and no third party reviewed them. The answer key comes from ReshapeX’s own answers, verified against manufacturer documentation, so ReshapeX is not scored on equal terms with the models it tested. Every question is published in the report so anyone can rerun the test.
Each model ran through its provider’s API with a pinned model ID (gpt-6-astra, claude-fable-5-1, gemini-3.8-flash): 1,500 calls on September 24, 2026. Gemini ran without web search because its provider’s terms require written permission to benchmark it with grounding. A model drafted each grade and Reshape’s engineers made the final call. The report lists 13 limitations.
The report prints all 100 questions and a SHA-256 fingerprint of the answer key. The key itself is withheld to keep the questions usable as a test and is available on request. Download the report at https://www.reshapex.com/iaai2026.
Trademarks
GPT-6 Astra, Claude, Gemini, Siemens, Festo, Rittal, ATI Industrial Automation and item are trademarks of their respective owners. None of them commissioned, reviewed, sponsored or endorsed this evaluation.
Download images: Here
Reshape Automation
Reshape Automation, Inc. builds AI agents that take on the tedious technical and commercial work of industrial sales and support. Its ReshapeX agents select and configure products, cross-reference parts and full bills of materials, prepare quotes, track orders, guide field troubleshooting and surface account insights for distributors, OEMs, system integrators and manufacturers. They work inside the customer’s ERP and CRM and run on websites, Outlook, Microsoft Teams, WhatsApp and voice. Each agent answers from a knowledge graph built from the manufacturer’s own catalogs and engineering knowledge, kept current as catalogs change, and Reshape’s forward-deployed engineers take every deployment into production alongside the customer. More at www.reshapex.com.
Contacts
Media contact
James Sugrue
Head of Marketing and Sales Operations, Reshape Automation
james@reshapeautomation.com

