Descripción
Anubis is the definitive benchmarking tool for local large language models on Mac. Built for Apple Silicon, it evaluates your LLMs with precision and transparency. Run a benchmark, then submit to the public leaderboard with one click (or turn on auto-submit in Settings). OPTIMIZED FOR APPLE SILICON A native SwiftUI app designed exclusively for macOS. It leverages Apple Silicon's unified memory architecture for accurate performance measurement with real-time hardware telemetry. No Electron andno web views - just a fast, responsive native experience. BENCHMARK MODULE Run comprehensive performance tests on any local model and watch real-time metrics as tokens stream in: tokens per second, time to first token, prefill (input) speed, average token latency, and total generation time. For reasoning models, thinking tokens are measured separately so your tok/s reflects real output - not chain-of-thought. Run a single pass, or repeat N times as a group and get the mean with a 95% confidence interval. Live hardware telemetry shows per-core CPU utilization (with efficiency/performance core breakdown), GPU usage, memory pressure, and thermal state - each chart expandable for a full drill-down. Every session is saved with full model metadata (quantization and format) so you can track improvements over time. Export any result as a CSV, or copy/save the results card as a shareable image. ARENA MODE Stop guessing which model is better. Arena puts two models head-to-head on the same prompt, in parallel or sequentially, with responses side by side and detailed timing breakdowns. Ideal for weighing quality versus speed, comparing quantization levels, or finding the right model for the job. MONITOR A live system metrics dashboard with real-time CPU, GPU, and memory charts. Launch the floating HUD for an always-on-top compact view while you work. Built-in stress tests let you saturate CPU cores, GPU compute (Metal), or memory bandwidth to validate your hardware. REPORTS Compare model performance across all your runs in one view - average tokens per second, time to first token, average token latency, and peak memory, side by side. Sort by any metric to find your fastest model instantly. VAULT + IN-APP MODEL DOWNLOADS Your central hub for model management. Browse models across every backend with full metadata - family, parameters, quantization, format, and size. Search the Ollama library and pull new models without leaving the app, with live, cancellable downloads. COMMUNITY LEADERBOARD Submit your results to the Anubis community leaderboard and see how your hardware and model configurations stack up against other Mac users. Submit manually or enable auto-submit in Settings. SUPPORTED BACKENDS - Apple Intelligence - on-device Foundation Models (macOS 26, supported hardware) - Ollama - full support for all Ollama models - LM Studio - rich metadata via the extended API - MLX-LM - Apple's MLX serving framework - vLLM, LocalAI, Docker Model Runner, Open WebUI - Any OpenAI-compatible server KEY FEATURES - Real-time token streaming with live performance metrics - Reasoning-aware tokens/sec, prefill speed, and an Ollama Thinking toggle - Repeat runs with mean +/- 95% confidence intervals and seed control - Per-core CPU + GPU charts with expandable drill-downs - Side-by-side model comparison in Arena mode - Live system monitor with floating HUD and CPU/GPU/memory stress tests - Browse and pull Ollama models in-app (cancellable) - Cross-run reports with sorting and filtering - Run history with filters, configurable limits, and multi-select delete - Community leaderboard with optional auto-submit - Shareable results card (PNG + clipboard) and CSV export - Hardware telemetry: CPU, GPU, memory, and thermal monitoring - 15 curated benchmark prompts across 5 categories - Configurable parameters (temperature, top-p, max tokens) - Reopens on your last-used backend and model
Novedades
Versión 3.9 · 23/8/2026Anubis 3.9 brings benchmark automation, a new backend, server-verified metrics, and the smoothest streaming yet. NEW: FLOWS - BENCHMARK AUTOMATION Stop babysitting benchmarks. Build a Flow once - a sequence of models, prompts, and settings - and run the whole thing with one click. Drag and drop steps in the visual editor, start from a built-in template, and walk away while Anubis works through the queue. Every Flow run is saved with a full report comparing every step, and Flows can be exported and imported as files to share with other Anubis users. NEW BACKEND: oMLX First-class support for oMLX, the open MLX serving stack. Anubis auto-detects an oMLX server, reads its server-measured performance stats, and adds management superpowers: browse and download models from Hugging Face inside the app, and load or eject models right from the Vault or the Benchmark toolbar. SERVER-VERIFIED METRICS When your backend measures its own performance, Anubis now uses it. LM Studio runs use LM Studio's native stats API for true server-side time-to-first-token and decode speed. llama.cpp timing blocks and oMLX usage stats are captured too. A new Anubis/Server toggle in Session Details lets you compare Anubis's measurements against the backend's own numbers - and a "server verified" badge shows when they agree. BUTTER-SMOOTH STREAMING A deep rework of live rendering fixes the stall-and-burst streaming some users saw on fast models: tokens now flow continuously, the tokens/sec chart tracks the real generation rate with no fake spikes, and long runs stay responsive from first token to last. ARENA, REPORTS, AND VAULT UPGRADES Arena: per-side thinking toggles for reasoning models, plus cleaner timing that excludes model load from time-to-first-token Reports: winner cards spotlight your fastest model, best TTFT, and most efficient setup at a glance Vault: see running models across every backend in one place, with load/eject controls and new backend filters QUALITY OF LIFE Model preparation options per run: use the model as-is, warm it up first, or force a cold load for worst-case numbers Model Load Time and Eval Duration now reported for more backends, not just Ollama Configurable stall timeout for slow or busy servers Assorted fixes and performance work throughout Enjoying Anubis? A quick rating on the Mac App Store genuinely helps a solo developer keep building. Thank you!
Información
- Vendedor
- John Taverna
- Categoría
- Developer Tools
- Versión
- 3.9
- Requiere
- iOS 15.0+
- Tamaño
- 10.6 MB
- Clasificación por edad
- 4+