Desert Ant launches 18 on-device AI models, claiming 300x real-time transcription on iPhone

Desert Ant launches 18 on-device AI models, claiming 300x real-time transcription on iPhone

  • Desert Ant Labs launched 18 on-device AI models, 12 stable and six beta, distributed through one SDK for Swift, Kotlin, and JavaScript; the company says they are free up to 100,000 monthly active devices.
  • Voz transcribes 10 minutes of audio in two seconds on an iPhone, according to Desert Ant, with word-level start and end timestamps; the company compares this with 4.7x faster performance than Whisper.
  • The 9 MB Clear model processes a five-minute laptop recording into what Desert Ant calls studio-quality audio in one second, while the 12 MB Redact model reports 88.8% PII detection across 27 languages.
  • Desert Ant says Detail 6, its video app planned for iOS 27, will replace its cloud APIs with local models, including a 284 MB Clips model that turns a 10-minute video into about a dozen clips in five seconds.
  • The company positions local inference as a privacy and sovereignty choice because customer data stays on the device, but Hacker News commenters noted that several initial models reuse or optimize existing open models and currently favor Apple platforms.

Hacker News opinions

I got excited about a fast transcription model, but Voz appears to be Parakeet v3 with new macOS and iOS-specific inference code.

We optimized Parakeet for the Apple Neural Engine with our own inference path, reaching about 300x real-time on iPhone 16 and 17. Our next Voz is trained from scratch and should be at least twice as fast, with Android and other platforms coming later.

I would get more use from a model that turns PDFs into a JSON schema, or generates titles and tags for posts. I am mainly interested in web apps.

We have a Schemer model coming for free-form text to structured JSON. Image-to-JSON is next on our list after that.

These models could improve the CMS work I do for clients, but most look iOS-only and the benchmarks use recent iPhones. I doubt a cheap VPS gets comparable speed.

Only a few models are iOS-first because of sequencing. We plan to make all models cross-platform in the coming weeks.

This looks like proprietary packaging and marketing over open models: Voz is Parakeet 0.6B v3, Clear is DeepFilterNet 3, and Ear is Whisper-tiny's language predictor.

With frontier LLM help, a recent Mac, and a recent iPhone, porting an existing model to the ANE is becoming a fairly straightforward hill-climbing exercise.

Why is Core ML the only target so far? Android has comparable options such as ML Kit.

Most models are planned for Android and web, though Voz, Clips, and Title take longer to port properly. The issue is mostly release order.

I like local specialized models even more than local LLMs, but I do not see the business model. If I already have the weights and run inference offline, why should the vendor keep earning from my customers?

You are conflating price with cost. A product can be worth paying for even when the vendor does not pay an inference cost per request.

Charging only after an app exceeds 100,000 monthly active devices seems unusually well aligned. A company at that scale can pay for the software it relies on.

Subscriptions pay for continuing updates as hardware and local-model tooling change, and a monthly expense is often easier for a startup to approve than a large upfront purchase.

The vendor owns the IP and sets the license terms. Downloading weights does not mean unrestricted use, and a free tier for 100,000 devices is generous.

A per-model-version purchase might fit local models better: buy the current weights, then decide whether a later update is worth paying for.

I tried the Clear demo and could not hear any difference between the raw and enhanced audio. Either my ear is not trained enough or the demo is broken.

DeepFilterNet 3 is small and fast but not especially good. MossFormer2 is better for commercially usable denoising, while Nvidia RE-USE can also remove reverb but has a non-commercial license.

On-device hate-speech triage sounds risky. What happens when moderation is automated locally?

I read it as a specialized model for uses such as game lobbies or parental filtering, where an indie developer needs moderation without paying large cloud-model bills.

AI
Solo Developer Trains 3.8B Model to 0.384 CORE for $998DeepSeek ships 552B V4.1-Flash, replaces V4-Pro with lower-cost multimodal modelRivian Prices Its Supervised Driving System Below Tesla While Building an AI Driver Around Temporal Object TrackingCognition says Devin-built GPU sieve factored RSA-260 for about $400,000GPT-5.5 reasoning prefills raise Qwen3.8 answer overlap by 18 points in a 45-problem testAnthropic's 2030 AI economy model ties rapid growth to weaker knowledge-worker jobsGPT-6 Astra Spurs Debate Over Looped Transformers and Hidden ReasoningOpenAI says GPT-5.6 Sol autonomously calibrated routine measurements on a six-qubit MIT chipOpenAI claims AI agents found a Navier-Stokes breakdown as credit dispute eruptsDesert Ant launches 18 on-device AI models, claiming 300x real-time transcription on iPhoneDeepSeek says V4.1 Flash will replace V4 Pro API traffic at lower pricesThoughtworks engineers turn a monorepo into an accidental agent blackboardOpenAI claims ChatGPT Images 2.5 cuts generation latency by up to 50%ICML paper finds LLM agents form new group biases from random feedbackInception releases Mercury 2.5 diffusion LLM, claiming 1,107 tokens/s and 260K contextTerence Tao Warns AI May Exhaust Mathematics' Supply of Fruitful Open ProblemsDeltafin streams 1.45 TB of Kimi K3 expert weights from four SSDs for 1 tok/s on an M5 Max MacBook ProDaVinci Resolve 21.1 adds Claude and ChatGPT Codex control for media, edits, and renderingGoogle DeepMind publishes AlphaGenome Atlas, predictions for 9 billion single-letter DNA variantsOpenAI says internal model proved finite-time Navier-Stokes singularityBuckmaster says LLM-assisted forced blowup work triggered dispute with OpenAIDan Luu tests 26 prompts and four skills for agentic Rust verificationMistral raises €3B at over €21B valuation for sovereign open-weight AIOpen-weight GLM 5.3-flash prompts a one-year warning on AI-driven vulnerability exploitationGoogle DeepMind unveils WeatherNext 3, an hourly global weather AI using live satellite dataPaper claims unpaired translation between embedding spaces exposes vector database privacy risksOpenAI says coding agents have reached research intern level, targets automated researcher by 2028OpenAI Chief Scientist Warns of Rapid Reasoning AI Progress and Calls for Broader Safety InterventionGPT-6 Astra placed blocks in bowls in 19 of 20 robot-arm trials, but matched Fable on puzzle insertionBryan Cantrill says detectable LLM prose drives readers away, points to Pangram as a spam-filter analoguePaper models LLM adoption as a contagion with tipping points into persistent dependenceCodeRabbit finds GPT-6 Astra catches 20% more cross-file bugs than GPT-5.6 SolAnthropic publishes a 29,511-module Lean 4 proof of Fermat's Last TheoremArtificial Analysis v4.2 adds private agentic and PDF tests, putting Claude Fable 5.1 firstOpenRouter lists GPT-6 Astra with 1M context, $10/$50 pricing and multi-provider routingEEBench measures AI circuit designs with SPICE, BOM cost, and tolerance-corner checksAnthropic says Claude formalized Fermat's Last Theorem in Lean in 11 daysSite Publishes Alleged Logs of OpenAI Agents Using Public Wikis to CoordinateStudy finds coding agents choose grep over LSP except when exhaustive references matterNolan Lawson says Claude now answers frontend performance questions that once sustained web-dev educationArmature's 16,893-run study finds Stripe, Neon and AWS dominate some coding-agent tool choicesGPT-6 Astra Scores 99.9% on ARC-AGI-3 With Persistent Reasoning StateClaude Code ports 72,758 lines of Amiga 68000 assembly to Godot, but fidelity remains manually judgedOpenAI says GPT-6 Astra hits 99.9% on ARC-AGI-3 and begins limited rolloutCerebras adds Qwen 3.8 27B public API endpoint rated at 1,500 tokens/sIFM releases six Apache-licensed K2 Horizon models with training checkpoints, data recipes, logs, and codeNvidia agrees to acquire Hugging Face for $12.9B, pledges to keep platform openNature study links LLM writing assistance to a 21 to 50% drop in writing-complexity varianceFable 5.1 builds a walkable Three.js model of San Francisco's Union SquareAisle reports six low-severity curl CVEs after Codex Security and Mythos returned zeroPerplexity cited 215,128 generated software-buying pages in tests across 380 categoriesMistral Vibe uses user data for training by default, while Vibe Enterprise is opted outGoogle launches Gemini 3.8 Flash at 3.7 pricing, adds restricted Cyber variantPaper approximates neural network representations with symbolic equations across LLM tasksOpenAI classifies Astra as its first Critical cyber-capable model and limits advanced accessWorld Labs unveils Atlas, a multimodal world model for 3D reconstruction, controllable video, and robotics simulationAnthropic launches Claude Fable 5.1 with lower cache-read pricing and a restricted Mythos 5.1 variantSmall transformer reaches 44% on ARC-AGI-1 for 67 cents after 1.5-hour RTX 5090 trainingApple's Enterprise AI Hardware Demand Outruns Mac Mini and Studio SupplyCornell guide traces how diffusion LLMs refine full text sequences in parallelOpenClaw 2.0 rebuilds setup, browser UI, and shared AI agent sessionsSimon Willison maps ChatGPT Work's cloud agent tools, shared Codex quota, and browser automationAI crawlers consume 14 CPU cores rendering kernel commits that could be clonedArtificial Analysis benchmarks sub-8GB local AI models on iPhone 17 Pro and Galaxy S26 UltravLLM 0.28 adds Kimi-K3 and DeepSeek V4 inference work, tiered KV cache offloadingTencent open-sources 770B-parameter Hy4 preview with 1M-plus-token contextDebian permits generative AI contributions, keeps contributors fully accountableSamsung puts MAC units in LPDDR5X banks for 614 GB/s in-memory AI computeLemmalog uses Datalog to retract stale LLM research conclusionsSamsung's LPDDR5X-PIM puts 16 compute blocks beside DRAM and reports 3.01x Llama 3.1 throughputStation’s autonomous AI agents report new results on five mathematical construction problemsOpenAI sets November 12 cutoff for Cursor after SpaceX acquisitionCohttp patch drew traversal probes within 10 minutes as agents turn bug hints into exploitsZ.ai releases GLM-5.3 open weights, claiming post-training gains in coding and cyber tasksEPA says off-grid data center power plants are outside Acid Rain ProgramUS Judge Blocks Pentagon's Anthropic Blacklisting as Illegal RetaliationTerminal-Bench-Science launches with 70 research workflows, Opus 5 reaches 30%Anthropic previews MHS, a device-driver standard for AI agents in labs and factoriesExperiential open-sources an LLM gateway that routes traffic and trains models from usageGoogle opens Gemini 3.5 Transcribe for real-time and recorded speech APIsGoogle opens Gemini Omni 1.1 Flash API with 40-second scene extension and 4K video upscalingBill Gates calls for democratic AI transition planning as job losses and data center impacts growCalvin French-Owen argues cheap, fast models make consumer and business AI economics viableMIT panel calls for AI-aware teaching and assessment, but leaves policy details to future workLAION releases 80M-video, 10M-hour research dataset for multimodal AI trainingBill Gates calls for AI taxes, protected jobs, and a global oversight bodyAmazon will permanently close Mechanical Turk on September 30, 2026Nvidia Is in Talks to Buy Hugging Face for More Than $13BOpenAI says internal GPT-5.6-scale model used Artifactory to bypass isolation and reach Hugging FaceAI Coding Risks Producing Developers Who Can Patch Systems Without Understanding ThemZ.ai releases 320B-parameter GLM-5.3-Flash, claiming near-Opus coding results at $0.045 per taskQwen opens Qwen3.8-Flash-Next, a 6B-active MoE previewing Qwen4 architectureZ.AI confirms Ox Alpha is a GLM-series model and plans to release its weightsResearcher says rooted Pixels can sign AI fakes as genuine C2PA camera capturesSemiAnalysis says OpenAI's Jalapeño inference ASIC beats Blackwell on throughput per MWApple updates Mac mini with M6, claims up to 4x faster AI performanceRL Trains Qwen 3.5 to Paint Editable Watercolours in JavaScriptApple's M5 Ultra Mac Studio reaches 512GB unified memory for local LLMs and four-node AI clustersApple puts a 2nm M6 in Mac mini and a 512GB-capable quad-die M5 Ultra in Mac StudioThomson Reuters spends $40M to train Thomson legal and tax LLM from open-weight models