Ollama

Rodar modelo aberto na própria máquina com um comando — e o release ritmado.

Ollama

2 fontes oficiais monitoradas · última atualização em 09/10/2026

Novidades recentes

40 atualizações

  1. Releases·oficial·

    v0.40.2

    Model upgrades Models downloaded with earlier versions of Ollama are upgraded in the background the first time you run them, for better performance and compatibility when running on llama.cpp. To make downgrading safe, Ollama keeps the original copy as a backup, so upgraded models are temporarily kept on disk. A future release will remove these backups automatically. To remove backed up models…

  2. Releases·oficial·

    v0.40.2-rc0: server: hide duplicate and downgrade guards from list (#18874)

    When a legacy GGUF which requires our llama.cpp patch set is loaded, we do a lazy conversion on disk to make it llama.cpp compatible. To support downgrades, we temporarily hold both the new and old GGUFs and create a shadow v1 manifest tag to prevent the downgraded server from deleting the new blobs. Showing these in the list output is confusing to users and API consumers. The legacy copies will…

  3. Releases·oficial·

    v0.40.1

    What's Changed server: proxy cloud usage and balance APIs by @drifkin in #18829 llama: fix clef head reads past 2GiB on windows by @Gigrise in #18777 cmd: remove account step from CLI onboarding by @hoyyeva in #18826 manifest: avoid symlinks on Windows by @dhiltgen in #18852 docs: fix 6 dead links in README community integrations list by @aniketkrs in #18814 docs: fix broken download links in app…

  4. Releases·oficial·

    v0.40.1-rc0

    mlx: drop carried metal residency patch now that it is upstream (#18854)

  5. Modelos·oficial·

    v0.40.0

    What's Changed Models run on MLX on Apple Silicon by default In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX. ollama pull qwen3.8 ollama run qwen3.8 Additional models include gemma4, qwen3.6 and qwen3.5 Decision models are now available on MLX as well: Nimble tev1 clef clef-flash MLX now has support for an embedding model:…

  6. Releases·oficial·

    v0.40.0-rc6: model: add multimodal embeddings (#18820)

    Implements the EmbeddingGemma2Model architecture on the MLX runner: 24-layer bidirectional text encoder with PLE, shared gemma4 vision/audio towers, mean-pool + L2 output. /api/embed accepts per-item media via input dicts.

  7. Releases·oficial·

    v0.40.0-rc5

    mlx: fix patch for recent mlx update (#18812)

  8. Releases·oficial·

    v0.40.0-rc4: MLX: version bump (#18720)

    MLX: version bump add scopes for unit tests to reduce memory usage

  9. Releases·oficial·

    v0.40.0-rc3: pull: allow RCs to pull matching min_version (#18790)

    Strip off pre-release from the version string so RCs can pull models for the same release.

  10. Releases·oficial·

    v0.40.0-rc2

    Merge remote-tracking branch 'upstream/main' into release_v0.40.0

  11. Releases·oficial·

    v0.40.0-rc1: mlx: match publisher tokenizer semantics (#18779)

    mlx: match publisher tokenizer semantics Honor pretokenizer stage order, split behavior, Unicode boundaries, added-token normalization, and ranked BPE merges. Handle empty added tokens and empty Metaspace input consistently. Add shared Go/Python reference cases using published tokenizers, pulling missing models directly and failing on errors, plus focused regressions for configuration precedence,…

  12. Releases·oficial·

    v0.35.1

    Clef decision models Ollama now supports Clef and Clef Flash, Cloudflare's new open-source decision models, through /v1/systemone. Clef (27B) and Clef Flash (9B) are multimodal: requests can now include images alongside the text state, shared by all questions and scored jointly with it. curl http://localhost:11434/v1/systemone -d '{ "model": "clef-flash", "state": "The user took this…

  13. Releases·oficial·

    v0.35.1-rc2

    ci: fix missing build context (#18742)

  14. Releases·oficial·

    v0.35.1-rc1

    models: add clef support (#18741)

  15. Releases·oficial·

    v0.35.0

    Decision models Ollama now supports decision models through /v1/systemone, based on TypeSafe’s Jev API. Decision models return choices, probabilities, and scores instead of text. Use them for tasks such as ticket triage, model routing, and content classification. Available models: Nimble from Bespoke Labs Tev1 from Together AI ollama pull nimble Send context and one or more questions: curl…

  16. Releases·oficial·

    v0.35.1-rc0: create: support explicit model capabilities (#18708)

    Add CAPABILITY declarations to Modelfiles and an additive capabilities field to create requests. Preserve declarations across GGUF and safetensors creation, inheritance, and Modelfile export. Require decision capability before scheduling System One requests instead of matching Qwen architecture/renderer metadata. Retain main's GGUF-only scoring restriction until the separate MLX runtime work…

  17. APIs·oficial·

    Ollama now supports Jev-style decision models

    Ollama now supports decision models, based on TypeSafe's Jev API for fast, typed decisions. Decision models can now be run at no cost with low latency. Based on text, decision models answer yes-or-no questions, choices and scores about it, with a probability for every option.

  18. Releases·oficial·

    v0.35.0-rc1

    mlx: bound pull stall retries and let the watchdog interrupt them (#1…

  19. Releases·oficial·

    v0.35.0-rc0

    feat: add System One scoring API (#18606)

  20. Modelos·oficial·

    v0.40.0

    What's Changed Models run on MLX on Apple Silicon by default In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX. ollama pull qwen3.8 ollama run qwen3.8 During the pre-release we will be testing and enabling additional models. Full Changelog: v0.34.4...v0.40.0-rc0

  21. Modelos·oficial·

    v0.34.4

    What's Changed Structured outputs on thinking models now apply in a single pass, making them faster and more reliable. Fixed intermittent "model not found" errors with a large local library Fixed the macOS app becoming unresponsive when checking if ChatGPT or Codex is running. Qwen 3.8 prompt processing is faster on Apple Silicon. Gemma 4 on Apple Silicon now picks the best image resolution per…

  22. Releases·oficial·

    v0.34.4-rc1: mlxrunner: Update XGrammar to 0.2.7 for structured outputs

    We pick up schema fixes for typed dictionary values and short arrays.

  23. Modelos·oficial·

    v0.34.4-rc0: mlx: speed up Qwen 3.8 prompt processing (#18550)

    mlx: speed up Qwen 3.8 prompt processing Use MLX's gated-delta kernel for long scans and fold dense MLP global scales into SwiGLU. address comments

  24. Releases·oficial·

    v0.34.3

    What's Changed GET /api/show now advertises each model's thinking controls and default: Available in the CLI with: ollama show gemma4 thinking levels false, true default true Available in the API with: curl http://localhost:11434/api/show -d '{"model": "glm-5.3-flash:cloud"}' { "thinking": { "values": ["low", "high", "max"], "default": "max" } } Also available on ollama.com directly for cloud…

  25. Releases·oficial·

    v0.34.3-rc1

    server: allow registry cross-host redirects among allowlisted hosts (…

  26. Releases·oficial·

    v0.34.3-rc0

    api: expose model thinking levels and defaults (#18473)

  27. Releases·oficial·

    v0.34.2

    What's Changed Added first-run setup when running ollama, with options to sign in or continue locally. Setup completion is shared with the desktop app on macOS and Windows. Added ollama://apps to open the desktop app’s Apps page directly on macOS and Windows. Fixed excessive memory growth during long generations with MLX speculative decoding. Updated llama.cpp. Full Changelog: v0.34.1...v0.34.2

  28. Releases·oficial·

    v0.34.2-rc3

    cli: add first-run onboarding shared with the desktop app (#18495)

  29. Releases·oficial·

    v0.34.2-rc2: mlxrunner: Release freed KV buffers during speculative decode

    The decode loop releases MLX's pool of freed buffers every 256 generated tokens, which is also how often the KV cache grows and drops its previous, smaller buffers. The check fires only when the token count lands exactly on a multiple of 256. Speculative decoding emits several tokens per round, so most rounds step over the boundary and the pool is never released. Each growth at a long context…

  30. Releases·oficial·

    v0.34.2-rc1: mlxrunner: lay out model by contract, checkpoint and construction

    model is one package with three jobs: the contract between the runner and the architectures, the opened checkpoint, and building nn layers from checkpoint tensors. Its files did not say which was which. base.go carried the folded package's name over the interfaces and the registry, root.go held the safetensors header scan next to Root, and quant.go mixed the nvfp4 global-scale helpers with quant…

  31. Releases·oficial·

    v0.34.2-rc0: llama.cpp: version bump b10969 (#18446)

    llama.cpp build changes resulted in duplicate symbols between libllama and libmtmd. This moves the compat patch into libllama with exported symbols.

  32. Releases·oficial·

    v0.34.1

    What's Changed MLX safetensors ollama create no longer experimental. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization. Improved MLX memory handling on Apple Silicon Runaway repeat token detection now requires 100 repeat tokens for reduced false positives (e.g. OCR) /api/tags is much faster on large model libraries (3.1 s → 294 ms cold in…

  33. Releases·oficial·breaking

    v0.34.1-rc2: API: Deprecate typical_p (#18448)

    Drop support for creating new models with typical_p parameters, while retaining support for existing GGUF models with the setting.

  34. Releases·oficial·

    v0.34.1-rc1

    mlx: add mlx patch to docker build context (#18440)

  35. Releases·oficial·

    v0.34.1-rc0: MLX: version bump (#18235)

    MLX: version bump mlx: support ModelOpt global scales in MoE models address comments address comments

  36. Releases·oficial·

    v0.34.0

    Use Ollama models in ChatGPT Desktop Ollama models can now be used directly in ChatGPT Desktop, so you can keep your existing workflow while running open models. Setup is available from the Ollama app on MacOS. This release also improves structured output performance on Apple Silicon, adds support for OpenAI-compatible client tool search and response compaction. Full Changelog: v0.33.3...v0.34.0

  37. Releases·oficial·

    v0.34.0-rc5

    openai: support standalone named function outputs (#18348)

  38. Releases·oficial·

    v0.34.0-rc4

    proxy: normalize namespaced commands in Full Access (#18331)

  39. Releases·oficial·

    v0.34.0-rc3

    openai: accept plaintext-labeled Codex agent messages (#18329)

  40. Releases·oficial·

    v0.34.0-rc2

    openai: finalize responses at the web search limit (#18328)

Fontes monitoradas

O radar lê estes endereços automaticamente. A ordem é a de prioridade: se a primeira falhar, as seguintes continuam entregando.

← Voltar ao radar