MikeTrendsTrends right now

search

local AI models

Trends

  1. 1
    Qwen 125B model claims fast inference on a single RTX 4090●Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/sYhnSportFootball91037 min ago

    A project called Strata claims to run Qwen's 125B-parameter Flash Next model on a single consumer RTX 4090 GPU at roughly 100 tokens per second. The claim drew heavy attention on Hacker News, where it ranked near the top of the site. If verified, it would represent a significant step for running large language models on everyday hardware rather than data-centre equipment.

  2. 2
    Turning off Apple Intelligence frees up disk space on macOS 27●Turn off Apple Intelligence on macOS 27 and get its disk space backYhnScienceSpace75519 min ago

    Mac users are discussing how to disable Apple Intelligence in macOS 27 and reclaim the disk space its models and caches occupy. A community script shared online automates the removal, and commenters are weighing whether the storage savings are worth losing the AI features, with some also raising privacy concerns about on-device models.

  3. 3
    Jevstiller distills Jev into a local model with disagreement bound●Show HN: Jevstiller – Distill Jev into a local model, with a disagreement boundYhnHealthNutrition6745 min ago

    A developer has launched Jevstiller, a tool that distills Jev into a local AI model, claiming a built-in disagreement bound that limits how far the distilled model can diverge from its source. The project was shared on Hacker News under the Show HN format, drawing moderate attention from the community. Details on the guarantee behind the disagreement bound are published on the project's site.

  4. 4
    Hacker turns iPhone into a second GPU for MacBook AI speedup●I made my iPhone a second GPU for my MacBook-Qwen 3.8 27B prefills 29–44% fasterYhnSportCricket381 h ago

    A developer has shown how to use an iPhone as a second GPU alongside a MacBook, reporting that prefill times for running the Qwen 3.8 27B model are 29 to 44 percent faster. The experiment exploits Apple's unified connectivity to share compute between devices for local AI inference. Commenters are debating how the setup works, its practicality, and what it means for running large models on consumer hardware.

  5. 5
    Developer builds offline AI planner that picks your best hour outdoors●Say what you want to do outside. Gemma, on your own computer, reads it, code picks the hour from an open forecast, and yMmastodonTechnologySoftware31 h ago

    A developer has built a small open-source tool that lets users type what they want to do outside; Google's Gemma model runs locally on the user's own computer, code picks the best hour from an open weather forecast, and the result comes back as a reminder and a pocket card that works without any signal. The project is being shared as part of an open-source development challenge.

Repos