AIO APEX

Anthropic builds in-house chip team to cut Claude inference costs and reduce Nvidia dependency

Tom's Hardware
Share:
Anthropic builds in-house chip team to cut Claude inference costs and reduce Nvidia dependency

Anthropic has confirmed it is assembling an in-house team of semiconductor engineers to design custom AI chips for its Claude models, joining Google, Microsoft, and Amazon in building proprietary silicon for AI inference workloads.

Why Anthropic Is Moving Into Custom Silicon

Despite maintaining partnerships with Google and Broadcom for TPU capacity, AWS Trainium processors, AMD MI450-series GPUs, and Nvidia hardware through Microsoft Azure, the economics of running Claude at production scale on third-party chips — particularly expensive Nvidia hardware — have pushed the company toward vertical integration. Anthropic describes the effort as a "multi-chip approach," focused on co-designing hardware and models together to maximize inference efficiency and reduce per-token costs as customer demand accelerates.

Samsung in Talks as Manufacturing Partner

Anthropic is actively recruiting engineers with "a proven track record of shipping finished semiconductor designs." Samsung Electronics is reported to be in exploratory discussions as the foundry partner, potentially using its advanced 2-nanometer manufacturing process. No formal deal has been announced and Anthropic has not disclosed a timeline for the first chip's readiness or whether it plans to self-manufacture any component.

Inference, Not Training — A Narrower Bet

Critically, Anthropic's custom chips are targeting inference workloads — the cost center for serving Claude to customers in real time — rather than training, which requires massive parallel GPU clusters. That is a narrower scope than what Google's TPUs initially tackled, and potentially a faster path to meaningful cost savings. Google's custom silicon took roughly four years of iteration before it meaningfully displaced Nvidia hardware for large-scale workloads; Anthropic's inference-first focus could compress that timeline.

The custom silicon effort complements, rather than replaces, Anthropic's existing hardware partnerships. The company has contracted approximately 3.5 gigawatts of Google and Broadcom TPU capacity coming online in 2027, over 300 megawatts from SpaceX's Colossus data center, and up to 2 gigawatts of AMD MI450-series chips. As reported by Tom's Hardware, Anthropic's in-house design team is in early recruiting and design phases, not yet at production stage.

The Race to Own the AI Stack

Anthropic's move mirrors the playbook of every major AI lab that has reached production scale. Google built TPUs. Amazon built Trainium and Inferentia. Microsoft is co-developing Maia with OpenAI. Meta is building its own Iris chip. The common thread: at sufficient scale, the margin lost to GPU rental economics is large enough to justify the capital and engineering cost of custom silicon. For Anthropic, which charges per Claude API token and competes on model quality and price, any material reduction in inference cost per token directly widens its competitive position.

Originally reported by Tom's Hardware. Read the original article for additional details.

View original source
Share:
Anthropic builds in-house chip team to cut Claude inference costs and reduce Nvidia dependency | AIO APEX