H100 Rents Dip 1.1% as AI Inference Costs Cool, Yet Demand Surges

Key Takeaways

J.P. Morgan reports synchronized drops in AI token and H100 rental prices in July 2026, signaling eased infrastructure tension. Despite lower costs, demand surges with high GPU utilization and rising open-weight model adoption.

Woofun AI reports that J.P. Morgan’s July edition of "Data Center Watch" reveals a synchronized cooling in AI inference call prices and high-end GPU short-term lease prices, indicating that the most tense phase of AI infrastructure is beginning to relax. This dual price decline suggests that the extreme scarcity driving previous market distortions is alleviating, allowing for a more normalized pricing structure across the AI compute stack. The data implies that while the immediate pressure on supply chains is easing, the underlying demand dynamics remain robust, challenging the notion that lower costs equate to waning interest in artificial intelligence applications.

In July 2026, the input Token price for mainstream models continued its downward trajectory, with sample data showing prices per million input Tokens generally ranging from $0.2 to $5. This range experienced an average monthly decline of approximately 3% during the month. It is crucial to contextualize these figures, as the output Token prices for high-performance closed-source models such as GPT-4o and Claude 3.5 Sonnet remain significantly higher than their input counterparts.

Furthermore, lightweight models like GPT-4o mini operate in a different pricing tier and cannot be directly compared to flagship offerings. Consequently, the reported price reductions primarily reflect trends in input Tokens, lightweight models, or specific hosted provider samples, rather than a uniform drop in the total cost of calling all mainstream models.

Despite these nuances, the directional trend offers tangible commercial benefits for AI application companies. The unit inference cost for common tasks such as question-answering, code generation, search enhancement, and customer service is decreasing. For years, the "the more you use, the more you lose" paradox has been a significant hurdle in the commercialization of large models. The downward pressure on Input Token unit prices is helping to bring high-frequency applications closer to an affordable cost range, thereby improving the economic viability of integrating AI into daily business operations.

This shift is particularly relevant for enterprises that rely on volume-driven AI services to maintain competitive advantages.

More importantly, this price reduction is occurring amidst a surge in inference volume, indicating that the market is not suffering from a lack of usage. From the first half of 2026 to July, the processing volume of LLM Tokens in the sample report increased several times compared to the beginning of the year, with some subcategories experiencing growth of over fourfold. Leading models from providers such as OpenAI, Anthropic, and Meta maintained steady high utilization rates between 70% and 80%.

Additionally, the widespread deployment of the Llama series of open-weight models significantly contributed to the overall increase in inference volume, demonstrating that the market is expanding rather than contracting.

The rapid growth of open-weight models emerged as another major theme in the July Token market. Data compiled by Woofun AI shows that in July, the usage of open-weight models increased by 26% month-on-month and surged by 63 times year-on-year compared to the same period last year.

Concurrently, the volume-weighted average price rose by 36% month-on-month, and Token expenditure increased by 71%. These metrics suggest that open-weight models are not merely competing on low prices but are gaining traction due to performance improvements that approach cutting-edge levels. As a result, the call prices, usage, and expenditure for some open-weight models are increasing simultaneously, reflecting a maturation of the open-weight ecosystem.

By Token usage, the top five models in July were MiMo v2.5, DeepSeek v4 Flash, GLM 5.2, DeepSeek v4 Pro, and MiniMax M3, collectively accounting for 50% of the Tokens' total supply. In contrast, by Token expenditure, the top five models were Claude Opus 4.8, Claude Opus 4.7, Fable 5, Kimi K3, and GPT-5.6 Sol, comprising 54% of the total expenditure. The entry of Kimi K3 into the top five expenditures signifies that open-weight models are moving into high-value territory traditionally dominated by closed-source models. Although closed-source models only represent 29% of the Tokens' total supply, they still contribute 80% of the expenditure, indicating that high-performance closed-source models continue to capture the majority of commercial value, even as open-weight alternatives gain ground.

The GPU rental market has experienced a similar structural differentiation. In July 2026, the average rental price of H100 in the non-hyperscale cloud provider market was $2.68 per GPU hour, marking a 1.1% month-on-month decrease. This represents the first month-on-month drop in H100 rental prices after seven consecutive months of increases. This decline is modest and falls well below the 15% to 25% range often cited in broader market discussions, and it cannot be characterized as a "20% drop in H100 rental prices." The report tracks the average monthly rental price calculated per GPU hour, rather than weekly rental prices per card, providing a more stable view of long-term trends. This slight decrease indicates a supply-side response, with more GPUs entering the cloud rental market and intensifying competition among cloud service providers and computational power platforms.

However, high-end resources remain expensive. The H100 price continues to be significantly higher than that of the A100, with overall GPU utilization rates maintained between 75% and 90%, and high-performance GPU utilization rates exceeding 80%. As long as utilization remains in this high range, the decline in rent appears more like a squeezing out of a portion of the scarcity premium rather than a sign of comprehensive oversupply in the computing power market. For AI companies, this change brings two distinct layers of impact. First, the cost of inference business is more likely to decrease, as inference requires a continuous, stable, low-latency supply of computing power. The decline in GPU rental prices will improve the cost structure for API service providers, AI search tools, code assistants, and enterprise AI products.

Second, the cost pressure of training large models is not easily alleviated by these rental price drops. High-end training relies not only on the number of GPUs but also on cluster interconnection, memory bandwidth, scheduling efficiency, and stable power supply. The drop in H100 rental prices does not translate to a proportional decrease in large-scale training costs, nor does it mean that all AI startups can access cluster resources of equal quality at low prices. The least thorough loosening of prices is seen in memory. From May to June 2026, sample data indicated a rapid increase in DRAM spot prices, with categories like DDR5 16Gb experiencing a cumulative increase of over 90%-100%, stabilizing only in July. In contrast, HBM, driven by AI accelerator demand, remains at high levels, with contract prices lagging behind spot prices by 1-2 quarters.

For high-end AI accelerators, the GPU chip itself is not the only bottleneck; HBM, advanced packaging, and data center delivery also impact real-world supply. Even as some GPU prices in the rental market fall, upstream memory and high-end cluster delivery will continue to limit the overall cost reduction rate. This is why the 'tightest moment of computing power begins to loosen' cannot be directly translated as the "end of computing power shortage."

The decline in rental market prices indicates that some short-term supply pressures have eased, but high-end training and large-scale inference clusters remain constrained by memory bandwidth, advanced packaging, and the pace of next-generation GPU shipments. The report reflects price and utilization changes in a July 2026 sample, serving as a snapshot of easing AI computing power in mid-2026 rather than a generalization for all cloud providers, regions, and long-term contracts.

The decrease in Token and GPU rental prices is positive for AI application companies but may squeeze profit margins for model service providers and cloud platforms, especially as open-weight models catch up and customer bargaining power increases. Variations in long-term contracts, regional differences, batch scale, and spot price calibers will lead to significant differences in actual prices obtained by different customers.

The most explicit signal is that by mid-2026, part of the AI infrastructure supply has caught up with demand, and unit calling and short-term costs are beginning to cool.

However, Token usage is still rapidly increasing, high-end GPU utilization remains high, output Token prices are higher, and HBM remains tight. Cost reductions are occurring first in segments easier to commoditize, while high-end clusters and upstream memory supply remain challenging. Welcome to join the official BlockBeats community: Telegram Subscription Group: https://t.me/theblockbeats, Telegram Discussion Group: https://t.me/BlockBeats_App, Official Twitter Account: https://twitter.com/BlockBeatsAsia.

Vote

Has the first H100 rent drop marked an AI compute turning point?

0 people voted

Comments

Me
Replying to @User
0/800

No comments yet.

Notifications

Sign in to view messages
View all messagesManage subscriptions