MikeTrendsTrends right now

search

AI inference hardware

Trends

  1. 1
    180B-parameter LLM runs locally on a laptop without a GPUโ—GPU ์—†์ด ์†Œ๋น„์ž์šฉ ๋…ธํŠธ๋ถ์—์„œ 180์–ต ํŒŒ๋ผ๋ฏธํ„ฐ LLM์„ ๊ตฌ๋™ํ•˜๋Š” POCKET-Darwin-180B. 4๋น„ํŠธ GGUF ์–‘์žํ™”๋กœ 360GBโ†’111GB ์••์ถ•, ์•ฝ $1,400 ํ•˜๋“œ์›จ์–ด๋กœ ๋กœ์ปฌ ์ถ”๋ก  ๊ฐ€๋Šฅ. # ai #MmastodonTechnologyAI358 min ago

    A project called POCKET-Darwin-180B is drawing attention for running a 180-billion-parameter language model on consumer hardware with no discrete GPU. Using 4-bit GGUF quantization, the model is compressed from roughly 360GB down to 111GB, enabling local inference on hardware costing about $1,400. Commenters in AI and open-source circles are highlighting it as a sign that frontier-scale models may soon run off the cloud.

  2. 2
    Power, Memory And Packaging Now Limit AI Chips, Not Transistorsโ–ผPower, Memory And Packaging, Not Transistors, Now Limit AI Chipsโœ‰newsTechnologySemiconductors54 min ago

    Industry analysts say the bottlenecks holding back AI chip performance are no longer transistor scaling. Power delivery, memory bandwidth and advanced packaging have become the key constraints, shifting how chipmakers like Nvidia, AMD and TSMC approach next-generation AI hardware design and investment.

  3. 3
    UC Berkeley and FuriosaAI Propose HBF for LLM Servingโ—HBF for High-Throughput LLM Serving (UC Berkeley, FuriosaAI)โœ‰newsTechnologyAI1 h ago

    Researchers at UC Berkeley, working with chipmaker FuriosaAI, have published work on HBF, a memory approach aimed at high-throughput serving of large language models. The piece, carried by Semiconductor Engineering, focuses on how new memory architectures could ease the bandwidth and cost bottlenecks that limit LLM inference at scale. The work is being followed by readers tracking hardware innovation for AI infrastructure.