search
CUDA
Trends
- 1Janus lets a single Go binary run GGUF models on any GPU●Show HN: Janus – Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia
A new open-source project called Janus is drawing attention on Hacker News. It is a Go binary that runs GGUF large language models through Vulkan graphics drivers, meaning it works on AMD, Intel and Nvidia GPUs without vendor-specific tooling. Commenters are discussing how it compares to existing inference tools and whether Vulkan can keep pace with CUDA-based performance.
- 2
Salvatore Sanfilippo, the programmer known as antirez who created Redis, has released ds4, an open-source engine for running DeepSeek 4 Flash and PRO models locally. The C-based engine targets Metal, CUDA and ROCm, meaning it can run on Apple silicon and AMD and Nvidia GPUs. The project appeared on GitHub and is drawing attention in the developer community.
- 3DeepSeek and Huawei open-source Ascend AI tools to rival Nvidia CUDA▼DeepSeek and Huawei release open-source Ascend AI programming tools to reduce reliance on Nvidia CUDA ecosystem — tools include compute and communication libraries, as well as Ascend support for TileLang
DeepSeek and Huawei have released open-source programming tools for Huawei's Ascend AI chips, including compute and communication libraries and Ascend support for TileLang. The move is aimed at reducing reliance on Nvidia's CUDA ecosystem for AI development, offering developers an alternative software stack and signaling growing momentum behind China's domestic AI hardware and software infrastructure.
- 4DeepSeek open-sources Huawei chip tools as CUDA alternative▼DeepSeek open-sources Huawei chip tools as a simpler alternative to CUDA
Chinese AI firm DeepSeek has released open-source software tools for Huawei chips, positioning them as a simpler alternative to Nvidia's widely used CUDA platform. The move would give developers a path to build AI applications on Chinese-made hardware amid continued US export restrictions on advanced chips.
- 5
Chinese AI startup DeepSeek has introduced software that allows its AI models to run on Huawei's Ascend chips, positioning the pairing as a domestic alternative to Nvidia's hardware. The move underscores China's push for technological self-sufficiency amid US export restrictions on advanced chips, and intensifies the competition between Huawei's ecosystem and Nvidia's dominant CUDA software platform.
- 6
Chinese AI firm DeepSeek and Huawei have opened up programming tools for Huawei's Ascend AI chips, making the technology accessible to outside developers. The move is seen as a step toward building a domestic AI software ecosystem around Chinese hardware, reducing reliance on Nvidia's CUDA stack amid ongoing US export restrictions. Details on the exact terms of the release remain limited.
- 7DeepSeek Open-Sources Full Ascend Infrastructure Stack▼DeepSeek Open-Sources Full Ascend Infrastructure Stack, Achieving "One-to-One Parity" with Nvidia Platform
Chinese AI firm DeepSeek has open-sourced its full infrastructure stack built on Huawei's Ascend chips, claiming one-to-one parity with its Nvidia-based platform. The move gives developers outside the CUDA ecosystem a complete, openly available software stack for training and running large AI models on domestic hardware. It is being read as a significant step toward viable alternatives to Nvidia amid ongoing US export restrictions on advanced chips to China.
- 8DeepSeek Open-Sources Ascend Versions of Key AI Libraries▼DeepSeek Open-Sources Ascend Versions of TileLang, DeepGEMM and DeepEP as Huawei Details SuperPoD Flex
DeepSeek has released open-source Ascend-compatible versions of its TileLang, DeepGEMM and DeepEP libraries, extending its AI software stack to Huawei's chip platform. The release came alongside Huawei detailing SuperPoD Flex, its flexible large-scale computing cluster architecture. The moves highlight deepening cooperation between China's leading AI lab and Huawei as both work to build alternatives to Nvidia's CUDA ecosystem.
- 9NVIDIA adopts Shibaura Institute computing method into CUDA●NVIDIA、芝浦工大発の計算手法をCUDAに採用。AI向けGPUの低精度演算器で「FP64相当」の計算を高速化 – ライブドアニュース https://www. yayafa.com/2901870/ # AgenticAi # AI #
NVIDIA has incorporated a computing method developed at Shibaura Institute of Technology into CUDA, its parallel computing platform. The technique uses low-precision arithmetic units on AI-oriented GPUs to accelerate calculations equivalent to FP64 double precision, promising faster scientific and AI workloads on consumer-grade hardware. The announcement, reported by Livedoor News, is drawing attention in Japanese tech and AI communities as a notable example of domestic Japanese research shaping GPU computing.
- 10DeepSeek open-sources toolkit for Huawei AI chips●DeepSeek has released an open-source software toolkit for Huawei's Ascend AI accelerators, including a programming langu
DeepSeek has released an open-source software toolkit for Huawei's Ascend AI accelerators, including a programming language called TileLang that rivals Nvidia's CUDA. The move supports Chinese AI development on domestic hardware and reduces reliance on Nvidia, whose chip sales to China face US export restrictions. Observers see it as a step toward an alternative software ecosystem for AI computing built around Chinese-made chips.
- 11Barracuda ransomware group claims Ecuador motoring club ANETA●🚨New ransom group blog post!🚨 Group name: Barracuda Post title: Automovil Club del Ecuador ANETA Location: 🇪🇨 EC Sector:
The ransomware group Barracuda has added Automovil Club del Ecuador ANETA to its leak site, listing the Ecuadorian motoring organisation in the services sector. The claim was flagged by a cyber threat intelligence feed that tracks new posts from ransomware gangs. No details have been released about the volume of data allegedly taken or any ransom demand.
- 12DeepSeek and Huawei Unveil Open-Source Toolkit to Challenge Nvidia CUDA▼DeepSeek and Huawei Unveil Open-Source Toolkit to Loosen Nvidia's CUDA Grip
DeepSeek and Huawei have jointly released an open-source toolkit aimed at reducing reliance on Nvidia's CUDA software ecosystem. The move pairs DeepSeek's AI research with Huawei's hardware push, offering developers an alternative stack for training and running AI models on Chinese chips. Industry watchers see it as a significant step in China's effort to build a self-sufficient AI computing stack amid US export restrictions on advanced Nvidia hardware.
- 13DeepSeek Teams Up With Huawei to Challenge Nvidia's Software Lead●DeepSeek Targets Nvidia's Software Moat With Huawei Partnership
DeepSeek is working with Huawei in an effort to break into Nvidia's dominance of AI software, where Nvidia's CUDA ecosystem has long been a key competitive barrier. The partnership signals a push to build a viable Chinese alternative spanning hardware and software for AI computing, intensifying the technology rivalry between US and Chinese chipmakers.
Repos
- lostmsu/TurboGPT Train a tiny GPT in under a minute (CUDA only)
- magnitudedev/magnitude Open source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on
- tensorflow/tensorflow An Open Source Machine Learning Framework for Everyone
- tile-ai/tilelang Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
- antirez/ds4 DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm