Login
Sign Up
Woofun AI reports that despite market panic framing the release of Kimi K3 as a 'DeepSeek Moment 2.0,' a consensus among Wall Street giants including UBS, Nomura, Bank of America Merrill Lynch, and Citigroup has emerged to reject this narrative. While Long Yue noted the initial fear that Chinese model advancements might curtail American compute spending, the prevailing institutional view posits that the sheer scale of Moonlight's new offering will instead act as a powerful accelerator for global infrastructure requirements. The divergence in sentiment stems from a fundamental reassessment of how open-source scale impacts the hardware value chain, shifting the focus from efficiency-driven cost cuts to volume-driven demand expansion.
Released on July 16 in Shanghai by Moonlight, Kimi K3 immediately established itself as a global benchmark with a staggering 2.8 trillion parameter count. The model achieved a score of 57 on the Artificial Analysis Intelligence Index, placing it third to fourth globally and on par with heavyweights like Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5. Performance metrics extended beyond general intelligence; on the Frontend Code Arena programming leaderboard developed by the University of California Berkeley, K3 secured the top position with a score of 1679. This achievement marked a historic milestone as the first open-source model to surpass all overseas closed-source competitors, including Claude Fable 5 and GPT-5.6 Sol, on an authoritative programming benchmark, signaling a shift in the competitive landscape.
The immediate market reaction on July 17 saw a significant decline in the U.S. semiconductor sector, echoing the volatility triggered by the release of DeepSeek R1 at the beginning of 2025. Investors reflexively questioned whether the rising capabilities of Chinese models would diminish the necessity for American AI companies to invest heavily in Nvidia chips, HBM, servers, and network equipment.
However, data tracked by the Wind Trading Desk indicates that major investment banks have issued reports contradicting this logic. The consensus from UBS, Nomura, Bank of America Merrill Lynch, and Citigroup is that Kimi K3 represents not a terminal point for compute demand, but a catalyst that will intensify the need for advanced hardware infrastructure across the globe.
Structurally, the impact of Kimi K3 differs fundamentally from the precedent set by DeepSeek R1, which highlighted efficiency gains. K3, released on July 16, 2026, with full model weights scheduled for opening on July 27, emphasizes massive scale and architectural complexity. The model features a 2.8 trillion parameter count, a 1M token context window, always-on inference capabilities, and native multimodal functions. Its architecture relies on Kimi Delta Attention, Attention Residuals, and Stable LatentMoE, activating 16 experts out of a pool of 896 for each token. Moonlight states that compared to Kimi K2, the overall scaling efficiency has improved by approximately 2.5 times, yet these advancements do not equate to a 'light asset' story. Instead, the combination of scale, context, and multimodal processing creates compounded pressure on inference, memory, network, and storage systems.
Pricing analysis compiled by Nomura reveals that Kimi K3 is positioned not as the lowest-cost option, but as a high-performance alternative approaching the capabilities of top-tier models at a reduced price point. The input cost is set at $3 per million tokens, with a cache hit input price of $0.30 and an output cost of $15, resulting in a total cost per task of approximately $0.94. This pricing structure places K3 below Claude Fable 5 at roughly $2.75 and Claude Opus 4.8 at $1.80, while remaining close to GPT-5.6 Sol's $1.04.
However, it remains significantly higher than GLM-5.2, which ranges from $0.32 to $0.47, and far above DeepSeek V4 Pro's $0.04. This positioning suggests that while K3 offers value, it is designed to compete on capability parity rather than pure cost minimization, ensuring that deployment requires substantial infrastructure investment.
Woofun AI data shows the argument against demand weakening was crystallized by four major investment banks, each highlighting different facets of the Jevons Paradox in the AI era. Duan Bing of Nomura's Asia-Pacific technology team argued that competition and innovation will persist as the industry approaches Artificial General Intelligence (AGI), driving continued investment from leading labs and hyperscalers to maintain competitive positions.
Peter Lee of Citigroup, in a July 19 report titled 'Another Jevons Paradox,' explained that efficiency improvements in AI models lead to greater consumption because more developers can afford to deploy applications and process more tokens. Lee emphasized that even with K3's efficiency, the expansion of KV cache usage driven by longer contexts will increase demand for server DDR5 and eSSD memory. Vivek Arya of Bank of America Merrill Lynch, writing on July 17, stated that the response from U.
S. labs like OpenAI, Anthropic, and Google will be 'not less computing power, but more,' as they strive to maintain differentiation through larger-scale training and heavier inference, especially given reports that Google's Gemini 3.5 Pro has faced delays. Timo Arcuri of UBS, in a July 20 report, reinforced that K3's 2.8 trillion parameters and 1 million token context window make it more reliant on HBM and storage than closed-source models, as the absolute demand for KV cache grows even after quantization.
Sector-specific beneficiaries are emerging clearly from this competitive dynamic, with storage identified as the most direct winner. UBS estimates that the cumulative free cash flow of the storage and memory sector will reach about 30% of market value by 2028, the highest proportion among all sub-sectors, with Micron (MU) alone accounting for 47%. Both Citigroup and Nomura maintain buy ratings on Samsung Electronics, citing extreme supply tightness in the global memory market.
Peter Lee noted that the large-scale deployment of Kimi K3 requires 'super-node' cluster configurations with more than 64 GPUs, directly driving demand for server DDR5 and enterprise-level solid-state drives. In the computing power infrastructure space, TSMC and Nvidia remain primary beneficiaries; Nvidia has stated that the inference performance-to-power ratio of modern MoE models on the GB300 NVL72 has improved by up to 25 times compared to the previous generation Hopper architecture.
Nomura reiterates buy ratings on TSMC, ASE, and MediaTek. Network infrastructure also stands to gain, as the super-node trend creates structural opportunities for optical module and chip suppliers like Zhongji Xuchuang and Suzhou Xuchuang, particularly in China where high-end chip export controls necessitate advanced architectures to bridge performance gaps. Cloud platforms such as Alibaba (BABA), GDS, and VNET are expected to benefit from ecological aggregation effects, leveraging their ability to host various cutting-edge open-source models.
Global penetration rates for Chinese AI models have accelerated dramatically, a data point often underestimated in broader narratives. Statistics from the open API gateway OpenRouter show that token usage by Chinese AI models accounted for less than 2% of global developer traffic a year ago but has now exceeded 45%. Data from Bank of America Merrill Lynch corroborates this acceleration, indicating that about 55% of U.S. enterprises have subscribed to AI models, platforms, or tools. Within this segment, Anthropic's enterprise adoption rate has reached 42%, while OpenAI stands at 40%.
Top AI consumers, representing the top 1% of enterprise users, have reached an AI spending level of $4,833 per employee per month. The market is diversifying, with Chinese open-source models like DeepSeek and Kimi K3 covering the economy and mid-to-high-end cost-performance markets, while top U.S. models focus on complex workloads such as scientific computing, maintaining technological and pricing premiums. Nomura judges that leading players on both sides will benefit provided they remain at the forefront of the technology curve.
The strategic shift initiated by Kimi K3 alters the rhythm of competition rather than the story of a single company. While the market initially compared K3 to DeepSeek R1, the latter forced a reassessment of training efficiency, whereas K3 demonstrates that open-source models can push scale, long context, agents, and multimodality to the cutting edge. This dynamic forces U.S. leading labs to continue investing heavily while allowing Chinese models to expand within the global developer ecosystem.
Closed-source top models retain their technological and price premiums, while open-source models cover a wider range of price points and deployment scenarios. In the short term, trading may fluctuate due to 'DeepSeek memory,' but in the medium term, as long as token usage grows and long contexts and agents spread, computing power, HBM, storage, network, and IDC will remain unavoidable cost items. Bank of America Merrill Lynch highlights that the hardware chain bears the training and inference pressure, ensuring sustained demand.
The final verdict from the financial community is that stronger open-source models are not the endpoint of AI infrastructure demand but may serve as the entry point for the next round of demand diffusion. Efficiency gains in models like Kimi K3 are expected to drive workloads higher, necessitating continued infrastructure construction.
However, Bank of America Merrill Lynch has identified a tail risk: if the speed of efficiency gains exceeds the growth of workloads, a pullback in infrastructure construction could occur. This caveat underscores that the logic of computing power demand growth relies on the synchronization of model affordability with significant expansion in usage, a condition that current market trends suggest is being met rather than threatened.