MikeTrendsTrends right now

search

local AI models

Trends

  1. 1
    Open-Source Edge Inference Engine Runs Large AI Models on Robots 10.7x Faster▼10.7x Faster: This Open-Source Edge-Side Inference Engine Enables Robot Bodies to Run Large Models Without Lag✉newsTechnologySoftware5 d ago

    A new open-source edge-side inference engine claims a 10.7x speedup, allowing robot hardware to run large AI models locally without lag. The technology targets real-time on-device inference for robotics, reducing reliance on cloud computing. Discussion is centered on its performance gains and what faster local inference could mean for embodied AI and robot deployments.

  2. 2
    AI Comes to Your Gaming PC●AI on Your Gaming PC https://hackaday.com/2026/10/04/ai-on-your-gaming-pc/ # AI # Gaming # HardwareMmastodonCultureGaming33 h ago

    Hackaday has published a new article looking at running AI directly on gaming PCs, focusing on the hardware side of the topic. The piece examines how consumer graphics cards and local machines can be used for AI workloads, a subject of growing interest among hobbyists and PC enthusiasts discussing local AI setups.

  3. 3
    Redis creator launches ds4 for running LLMs locally●From the creator of Redis; run LLM locally with ds4YhnTechnologyAI3521 h ago

    Salvatore Sanfilippo, the creator of Redis, has introduced ds4, a tool for running large language models on local machines. The project, hosted under the Dwarfstar name, is drawing attention among developers interested in self-hosted AI. Commenters are discussing its approach to local inference and what the involvement of a well-known open source figure means for the project's prospects.

  4. 4
    AI debate: on-device compute or data centers?●🤖 Will the AI compute crunch be solved on-device or in data centers? I build iOS apps and I'm pushing as much as possiblMmastodonTechnologyAI25 d ago

    An iOS developer is weighing whether the growing demand for AI computing power will ultimately be met on devices or in data centers, saying they push as much processing on-device as possible for privacy and cost reasons. They note Apple is betting on on-device AI, but argue frontier models keep getting bigger, and are asking where others think the balance will land.

  5. 5

    Salvatore Sanfilippo, the programmer known as antirez who created Redis, has released ds4, a local inference engine for running DeepSeek 4 Flash and PRO models. The C-based engine targets Apple Metal, CUDA and ROCm, letting users run the DeepSeek models on their own hardware across NVIDIA, AMD and Apple Silicon GPUs. The project is drawing attention in open-source AI circles.

  6. 6
    Qwen 3.8 Flash Next 125B claimed to run fast on RTX 4090▼Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/sYhnSportFootball15812 min ago

    A project called Strata, shared on GitHub, claims it can run Qwen's 3.8 Flash Next 125-billion-parameter model on a single consumer RTX 4090 GPU at roughly 100 tokens per second. If verified, that would make a very large language model practical on high-end home hardware without a data center. The claim is drawing attention among developers interested in local AI inference, though independent confirmation of the speed figures has not been established.

  7. 7
    Using an iPhone as a second GPU speeds up local AI models▼I made my iPhone a second GPU for my MacBook-Qwen 3.8 27B prefills 29–44% fasterYhnSportCricket3133 min ago

    A developer reports hooking up an iPhone to a MacBook as an extra compute device, accelerating prefill times for the Qwen 3.8 27B AI model by 29-44%. The trick taps the phone's Apple Silicon GPU alongside the laptop's own, drawing interest from people running large language models locally on consumer hardware.

  8. 8
    New tool runs pretrained classifiers locally without a GPU▼Show HN: Local pretrained classifiers, GPU not neededYhnWorldElections81 h ago

    A developer has released Jeffy, an open-source tool for running pretrained machine learning classifiers on local hardware without requiring a GPU. The project, shared on Hacker News, is aimed at making lightweight classification accessible to users with ordinary computers. Early engagement is modest, with the community beginning to evaluate its practical usefulness.

  9. 9
    Local AI Models Now Run Smoothly on Consumer Gaming Hardware●Local AI Models Run Smoothly on Gaming PCs and Laptops𝕏xSE27622 h ago

    Locally run AI models are reportedly operating smoothly on ordinary gaming PCs and laptops, without cloud servers or subscriptions. The discussion centers on how modern GPUs and increasing memory in consumer machines are enough to handle open-source language models at home. Commenters highlight growing interest in private, offline AI use and note that hardware once bought mainly for games is now doubling as a capable local AI workstation.

  10. 10
    180B-parameter LLM runs locally on a laptop without a GPU●GPU 없이 소비자용 노트북에서 180억 파라미터 LLM을 구동하는 POCKET-Darwin-180B. 4비트 GGUF 양자화로 360GB→111GB 압축, 약 $1,400 하드웨어로 로컬 추론 가능. # ai #MmastodonTechnologyAI31 d ago

    A project called POCKET-Darwin-180B is drawing attention for running a 180-billion-parameter language model on consumer hardware with no discrete GPU. Using 4-bit GGUF quantization, the model is compressed from roughly 360GB down to 111GB, enabling local inference on hardware costing about $1,400. Commenters in AI and open-source circles are highlighting it as a sign that frontier-scale models may soon run off the cloud.

  11. 11
    Pink Slime partisan content is seeping into AI chatbots●'Pink Slime' Is Infecting AI Chatbots Ahead of the MidtermsYhnWorldPolitics107 h ago

    Politico reports that 'pink slime' — networks of partisan local-news sites that mimic legitimate journalism — is now shaping the output of AI chatbots ahead of the US midterm elections. Because chatbots often cite or summarize online articles, low-credibility political content can be laundered into seemingly neutral answers, raising fresh concerns about election misinformation and how AI models source their information.

  12. 12
    Developers Turn to Mac Minis for Running AI Models●Why Developers Are Running AI Models on Mac Minis Instead of Nvidia GPUs✉newsBusinessStartups6 d ago

    Developers are increasingly running AI models on Apple's Mac Mini instead of relying on Nvidia GPUs, according to a Fortune report. The shift is being attributed to the Mac Mini's lower cost and power efficiency, with Apple silicon offering competitive performance for local AI workloads. The trend highlights a challenge to Nvidia's dominance in AI hardware as smaller teams look for cheaper ways to build and test AI applications.

  13. 13
    PewDiePie Says OpenAI Banned Him Twice While Building His Own AI▼PewDiePie Says OpenAI Banned Him Twice While He Built Ajax, His Own Local AI Model✉newsTechnologyGadgets1 h ago

    PewDiePie, the YouTuber, says OpenAI banned him twice while he was developing Ajax, a local AI model he built himself. The claim, reported by Gadget Review, highlights his move toward running AI on his own hardware rather than relying on mainstream services. His years of friction with OpenAI and his shift into self-hosted tech are drawing attention from both fans and the AI community.

  14. 14
    Morocco Releases Open-Source Darija AI Tools With Mistral▼Morocco Releases First Open-Source Darija AI Tools From Mistral Partnership✉newsTechnologySoftware6 h ago

    Morocco has released its first open-source AI tools for Darija, the Moroccan Arabic dialect, developed in partnership with French AI company Mistral. The release marks a step toward building AI systems that understand the local language, which is underrepresented in mainstream models. Observers see it as part of Morocco's push to develop sovereign AI capabilities.

  15. 15
    Codex plugins can now be used inside Pi coding agent●Show HN: Use all Codex Plugins inside Pi I just realized that codex now exposes local server endpoints for all plugins wMmastodonBusinessStartups31 d ago

    A developer has discovered that Codex exposes local server endpoints for all of its plugins without extra authentication, meaning those plugins can be used from any other model or agent harness. A new Pi install package lets users connect to all Codex plugins with a single auth setup. Developer communities are discussing what this means for interoperability between AI coding tools and whether open local endpoints could raise security questions.

  16. 16
    Multi-Token Prediction Boosts RTX 3090 LLM Speed▼Originally published on my blog. Enabling MTP on this RTX 3090 raised generation throughput from... # ai # llm # programMmastodonTechnologySoftware51 d ago

    A developer reports enabling multi-token prediction (MTP) on an RTX 3090 graphics card raised local LLM generation throughput, while questioning whether the speedup affects coding quality. The write-up, originally published on a personal blog, has drawn attention from AI and open-source software communities interested in getting more performance from consumer GPUs for running large language models locally.

  17. 17
    Writer ditches Grammarly for a local AI model●I replaced Grammarly with a local LLM, and none of my writing leaves my laptop anymore✉newsTechnologyAI1 h ago

    An XDA Developers article describes replacing Grammarly with a locally run large language model so that all writing stays on the author's laptop. The piece highlights growing interest in offline AI tools that handle grammar and editing without sending text to cloud services, appealing to privacy-conscious writers.

  18. 18
    Open source tool lets you run Jev locally▼Open source tool distills Jev so you can run it locally✉newsTechnologySoftware3 d ago

    The Register reports on a new open source tool that distills Jev, making it possible to run it on local hardware rather than in the cloud. Distillation shrinks a model so it can run on ordinary machines, lowering cost and keeping data private. The piece describes the tool and what it means for developers wanting offline use.

  19. 19
    Philadelphia Inquirer launches AI tool Scrape for hyperlocal news●The Philadelphia Inquirer built Scrape, an AI tool to surface hyperlocal news https://www.lenfestinstitute.org/solutionsMmastodonTechnology31 d ago

    The Philadelphia Inquirer has built Scrape, an artificial intelligence tool designed to surface hyperlocal news that might otherwise go unreported. The project is highlighted by the Lenfest Institute, which supports the paper and promotes it as a model for local journalism. Observers in tech circles are discussing whether AI can help struggling local outlets cover neighborhood-level stories at scale.

  20. 20
    Telegram bot reads bills locally with Gemma and nagging reminders●A Telegram bot that reads bills with local Gemma and keeps reminding until you pay, with durable reminders on Temporal aMmastodonTechnologySoftware216 h ago

    A developer has built an open-source Telegram bot that uses Google's Gemma model running locally to read and understand bills, then sends persistent reminders through Temporal's durable workflow system until the bill is paid. Sentry is used for agent tracing without exposing bill data. The project is part of a weekend coding challenge and is being shared openly with the developer community.

  21. 21
    Apple Mac Studio with M5 Ultra runs frontier AI models locally▼Apple Mac Studio (M5 Ultra) Review: Unlimited Power The Mac Studio can run frontier-level AI language models locally. ItMmastodonTechnology22 d ago

    A new review of Apple's Mac Studio with the M5 Ultra chip says the desktop can run frontier-level AI language models locally, calling it a preview of what's to come. The Wired verdict, summarised as 'unlimited power', is drawing attention for suggesting high-end local hardware can now handle AI workloads previously reserved for cloud data centres.

  22. 22
    AI model sizes mapped from 100KB to 2.5TB●🤖 Everyone is obsessed with trillion-parameter models, so I mapped out the entire AI spectrum from 100KB to 2.5TB (and wMmastodonTechnologyAI15 h ago

    A new overview charts the full range of AI model sizes, from tiny 100KB models running locally to trillion-parameter giants weighing 2.5TB, alongside what each size costs to run. It argues the industry conversation is fixated on massive datacenter systems and hourly H100 rentals, while the smaller end of the spectrum goes largely unexamined.

  23. 23
    The Exercise Coach Brings AI-Assisted Workouts to Jacksonville▼AI-assisted fitness and workouts now available at The Exercise Coach in Jacksonville✉newsHealthFitness1 d ago

    The Exercise Coach, a fitness studio franchise, has introduced AI-assisted training at its Jacksonville location. The technology personalizes strength-training workouts, adjusting exercises to each client's ability and progress. The rollout highlights a broader trend of artificial intelligence entering the fitness industry, with local media noting the studio's machine-guided, time-efficient workout model now paired with AI-driven coaching tools.

  24. 24
    Federal judge calls Flock surveillance system indiscriminate mass surveillance●Privacy & security, Sun, Oct 4: • Federal judge calls Flock 'indiscriminate mass surveillance' https:// techcrunch.com/2MmastodonTechnologyCybersecurity22 h ago

    A federal judge has sharply criticized Flock, the automated license plate reader company, describing its camera network as 'indiscriminate mass surveillance.' The ruling adds to mounting legal scrutiny of Flock's partnerships with local police departments across the United States. Privacy advocates are amplifying the decision alongside other security concerns, including Anthropic asking Claude users to share voice recordings for AI model training.

  25. 25
    Framework Desktop with 192GB memory opens pre-orders●192GB Framework Desktop open for pre-orderYhnLifeHome & Garden1616 h ago

    Framework has opened pre-orders for its Desktop DIY configuration built around AMD's AI Max 400 platform, with configurations supporting up to 192GB of unified memory. The unusual memory capacity, rare in a compact desktop, is drawing attention from developers and enthusiasts running local AI workloads, who see it as a flexible alternative to traditional mini PCs.

  26. 26
    Citi: Open-Source AI Model Threat to Frontier Revenues Has Peaked▼[Major Bank] Citi: Impact of Open-Source Weight Models on Frontier Revenues Has Passed Local Peak; Capability Gap Widens Again✉newsTechnologySoftware1 d ago

    Citigroup analysts argue that the revenue pressure open-source weight models once placed on frontier AI developers has passed its local peak, and that the capability gap between leading closed models and open alternatives is widening again. The note, circulating on financial news feeds, suggests investors may reprice AI lab revenues as closed frontier models regain their technical lead.

  27. 27
    Ten-minute seated pose sketch shared by German drawing studio●Sitzende, Fineliner und Fasermaler auf Papier, Pose zehn Minuten. Modell: Jana # aktzeichnung # schnellestudien # art #MmastodonCultureArt420 h ago

    Atelier am Kirschgarten, a drawing studio in Germany, shared a quick life-drawing study of a seated figure by model Jana, drawn on paper with fineliners and felt pens in a ten-minute pose. The piece is part of the studio's regular figure-drawing sessions and quick-sketch practice, shared with the online art community alongside tags for traditional, human-made artwork.

  28. 28
    Apple overhauls macOS Full Disk Access to curb AI agents▼Apple is updating macOS security by overhauling its Full Disk Access permission model in direct response to AI agents. TMmastodonTechnology21 d ago

    Apple is revamping macOS's Full Disk Access permission model in response to the rise of AI agents. The new approach is designed to stop agentic apps from demanding broad, persistent access to sensitive data such as local files, Mail stores, iMessage databases and browser histories. Observers see it as a direct acknowledgement that AI software is reshaping what desktop security rules must protect against.

  29. 29
    Redis creator launches ds4 for running LLMs locally●From the creator of Redis; run LLM locally with ds4 Article URL: https:// dwarfstar.sh/ Comments URL: https:// news.ycomMmastodonTechnology21 d ago

    A new tool called ds4, promoted as coming from the creator of Redis, lets users run large language models on their own machines. The project is being shared on developer forums, where early readers are weighing its promise of private, local AI inference. Details on features and licensing remain thin, and discussion is just beginning.

  30. 30
    Developer Breaks Down llama.cpp Configuration for Qwen 3.8B●Understanding My llama.cpp Qwen 3.8 Configuration I've been tuning llama.cpp for local AI development, and the command lMmastodonTechnologyAI21 d ago

    A developer has published a parameter-by-parameter walkthrough of their llama.cpp setup for running the Qwen 3 8B model locally, explaining what each command-line flag does and how the options are tuned for maximum performance on their hardware. The guide is aimed at people running AI models on their own machines, where cryptic command-line options often make local inference setups hard to understand and reproduce.

  31. 31
    UC Santa Cruz's Adam Smith on local small language models●Adam Smith from UC Santa Cruz joins us to discuss local Small Language Models (SLMs) and building open, autonomous toolsMmastodonTechnologyAI22 d ago

    Adam Smith of UC Santa Cruz is discussing the case for running small language models locally rather than relying on large cloud providers. He presents BayLeaf AI, described as a counterplatform, along with the concept of "transagency" — a human-agent collaboration model he likens to the relationship between a driver and a car. The conversation also covers context distillation and practical approaches to building open, autonomous AI tools that users control themselves.

  32. 32

    The GLM 5.3 Flash model is reportedly capable of running at frontier-level performance on a pair of Nvidia DGX Spark desktop systems, according to the claim drawing attention online. The setup suggests advanced AI inference can now be achieved on compact, relatively affordable local hardware rather than large data centre clusters. Commenters are discussing the implications for accessible high-end AI.

  33. 33
    PewDiePie builds local AJAX AI model after OpenAI ban●Pewdiepie builds smaller, local AJAX AI model after OpenAI ban controversy✉newsTechnologyAI17 h ago

    YouTuber PewDiePie has built a smaller, locally run AI model called AJAX, following a controversy involving a ban from OpenAI. The model runs on local hardware rather than cloud services, reflecting a shift toward self-hosted AI tools. The story is being covered by tech outlets, with attention focused on both the OpenAI dispute and the move to independent, offline AI development.

  34. 34
    NVIDIA DGX Spark 64GB expands local AI development options●NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI✉newsTechnologySoftware1 d ago

    NVIDIA has announced the DGX Spark with 64GB of memory, a compact AI development system aimed at giving developers more ways to build and scale AI applications locally. The company says the machine lets developers prototype, fine-tune and run AI models on their desktop without relying on cloud infrastructure.

  35. 35
    PewDiePie Says OpenAI Banned Him Twice Over Local AI Model●PewDiePie Claims OpenAI Banned Him Twice Over Local AI Model𝕏xSE572 d ago

    YouTuber PewDiePie, real name Felix Kjellberg, claims OpenAI banned his account twice, which he attributes to his use of local AI models on his own hardware. The claim, made publicly by the creator himself, has drawn attention from tech communities debating platform moderation and the push toward self-hosted AI. OpenAI has not publicly commented on the alleged bans.

  36. 36
    Local AI decision model Bespoke Nimble draws experimenter interest●I’ve been experimenting with Bespoke Nimble, a local decision model running through Ollama. It takes evidence, a questioMmastodonTechnologyAI11 d ago

    A developer is testing Bespoke Nimble, a small decision model run locally through Ollama. The model takes evidence, a question and a set of allowed answers at request time, meaning the same model can handle many classification tasks without retraining. The author is comparing it with another model called Jev in a write-up, and interest centres on whether compact local models can replace task-specific trained classifiers.

  37. 37
    Anthropic proposes opt-out AI training rules for Australian content●TL;DR: AI company Anthropic calls for an opt-out model for Australian content to train its models, while ABC and SBS demMmastodonTechnology32 d ago

    Anthropic has told an Australian review that AI firms should be able to use locally published content for training unless creators opt out. The proposal puts the company at odds with Australian broadcasters ABC and SBS, who are demanding strict regulations to protect journalism and ensure media organisations are fairly compensated when their work trains AI models.

  38. 38
    TensorFold claims up to 3x faster LLM inference on Mac and DGX Spark●シタン先生もpythonについて話していました Mac・DGX SparkでLLM推論を最大3倍高速化する「TensorFold」の概要|npaka https:// note.com/npaka/n/n3d3e09549bdd # AppMmastodonWorld32 d ago

    A new tool called TensorFold is being described as able to speed up LLM inference by up to three times on Apple Macs and Nvidia's DGX Spark hardware. A Japanese-language explainer by npaka on Note is circulating, and comments reference discussions of Python in relation to the tool. The claim is drawing attention among AI developers interested in running large language models locally.

  39. 39
    Bilibili Open-Sources Translation Model Family Covering 150 Languages▼Bilibili Open-Sources Index-Translate, a Qwen3.5-Based Translation Model Family for 150 Languages✉newsTechnologySoftware3 d ago

    Bilibili has open-sourced Index-Translate, a family of translation models built on Alibaba's Qwen3.5 that supports 150 languages. The release puts a large multilingual translation capability into open weights, letting developers run and fine-tune it themselves. The move adds to a growing wave of Chinese tech firms releasing open-source AI models and could draw interest from localization and machine translation developers.

  40. 40
    Using a local LLM to clean up a full hard drive●I gave my local LLM a nearly-full SSD and told it to find everything I could safely delete✉newsTechnologyAI3 d ago

    A tech writer describes running a locally hosted large language model on a nearly full SSD, asking it to identify files that could be safely deleted. The piece highlights a practical, off-cloud use of local AI: letting the model scan the drive and suggest disk cleanup targets, reflecting growing interest in running LLMs directly on personal hardware for everyday tasks.

Repos