Together AI

Inferência e fine-tuning de modelos abertos, com anúncios frequentes de preço.

Together API

1 fonte oficial monitorada · última atualização em 06/10/2026

Novidades recentes

40 atualizações

  1. Outros·oficial·

    Expanding our enterprise inference capacity with IBM Cloud and NVIDIA

    Enterprises can now run open models at production scale on a dedicated B300 inference cluster, built by Together AI, IBM Cloud, and NVIDIA

  2. Outros·oficial·

    Together Link: open models in the harness you already use. Start with one command today.

    Together Link brings frontier open models like GLM 5.3 and Kimi K3 into the coding agent your team already uses, cutting model spend by over 50%.

  3. Modelos·oficial·

    How to train your own Jev for $17

    We just launched our own Jev-like classifier, together/Tev1-4B-experimental, on top of Qwen3.5 4B on Together’s serverless platform. In this blog post we’ll show you how to fine-tune your own version!

  4. Outros·oficial·

    Canary rollouts: upgrade models in production without downtime

    A hard model swap exposes every user at once, and rolling back means cold-starting the old deployment under pressure. Here's how staged traffic ramps, metric gates, and automatic rollback work on dedicated inference.

  5. Outros·oficial·

    How a global fintech scaled coding agent traffic with Dedicated Model Inference

    Inside a global bank's shift to self-serve dedicated inference: how Together's DMI gave engineering teams direct control over scaling, models, and testing.

  6. Outros·oficial·

    Migrating from closed to open source models, Together

    Moving from closed to open source models can take weeks, not years. A five-stage playbook: discover, evaluate, adapt, decide, and production.

  7. Preços·oficial·

    Together AI expands fine-tuning service with more models, live metrics, and finer controls

    Together Fine-Tuning adds the latest open-weight models, live experiment tracking, Expert LoRA, early stopping, tokenized dataset previews, pre-flight validation, and lower training prices on selected models.

  8. Preços·oficial·

    Introducing preemptible compute: the same compute, half the price

    Together GPU Clusters now supports preemptible compute: the same GPU capacity at a flat 50% of the on-demand rate, with a five-minute drain window.

  9. Outros·oficial·

    To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

    We ported ThunderKittens to NVIDIA's Vera Rubin NVL72 and rebuilt our NVFP4 GEMM around the new hardware, taking it from 42% of roofline to over 22 PFLOPS — competitive with cuBLAS and CuTe DSL. Here is what changed in the ISA and how we used it.

  10. Outros·oficial·

    The Open Source AI Stack

    A deep dive into the open model AI stack — model, inference, gateways and routers, harness, and tools — and how keeping each layer independent lets you swap in a new open model in minutes instead of rebuilding your workflow.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  11. Outros·oficial·

    GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing

    We ran 900 DeepSWE rollouts on GLM-5.3 and GLM-5.3 Flash. Flash gives up 5.6 points of pass@1 at 17x lower cost, and only 2.6 points at pass@4.

  12. Modelos·oficial·

    GLM-5.3 vs. GPT-5.6 Sol no DeepSWE: custo, código e roteamento

    Foram executadas 904 rodadas do DeepSWE com GLM-5.3 e GPT-5.6 Sol. O Sol lidera no pass@1 por 3,7 pontos; o GLM-5.3 vence no pass@4 pela metade do custo, e uma cascata começando pelo GLM chega a 85,9%.

    traduzido por IA · claude-opus-5

    original: “GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing”

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  13. Features·oficial·

    GLM-5.3 vs. Claude Fable 5 no DeepSWE: custo, código e roteamento

    Foram executadas 904 rodadas do DeepSWE com GLM-5.3 e Claude Fable 5. Empate no pass@1, mas o GLM-5.3 vence no pass@4 e custa 5,4× menos: US$ 3,99 por rodada contra US$ 21,63.

    traduzido por IA · claude-opus-5

    original: “GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing”

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  14. Modelos·oficial·

    DeepSeek V4 Pro 0813 vs. GPT-5.6 Sol no DeepSWE: custo, código e roteamento

    Foram executadas 904 rodadas do DeepSWE com DeepSeek V4 Pro 0813 e GPT-5.6 Sol. O Sol lidera no pass@1 por 10 pontos, mas custa 35× mais; o Pro vence no pass@4, e uma cascata começando pelo Pro chega a 83,0%.

    traduzido por IA · claude-opus-5

    original: “DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing”

  15. Outros·oficial·

    DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

    We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  16. APIs·oficial·

    Teste A/B de modelos em produção

    Tráfego sombra prova que um candidato é operacionalmente sólido, mas não diz se os usuários gostam mais dele. Faça a divisão no endpoint, e não dentro do seu código.

    traduzido por IA · claude-opus-5

    original: “A/B test models in production”

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  17. Modelos·oficial·

    DeepSeek-V4 Flash 0731 vs. GPT-5.6 Luna no DeepSWE: custo e código

    Foram executadas 900 rodadas do DeepSWE com DeepSeek-V4 Flash e GPT-5.6 Luna. O Luna lidera no pass@1 por 14 pontos; o DeepSeek entrega 4,8× mais soluções por dólar.

    traduzido por IA · claude-opus-5

    original: “DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding”

    data informada pela fonte; encontrada pelo radar em 10 de set. de 2026

  18. APIs·oficial·

    Kimi K3: the complete developer guide

    Kimi K3 is the first open 3T-class model. See how it benchmarks, what it costs, and how to call it on the Together AI API, with copy-paste code examples.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  19. Outros·oficial·

    Autoscaling endpoints for LLM inference

    GPU utilization can read healthy while your queue backs up, and a new replica takes minutes to warm. Here's how to pick autoscaling metrics, tune scale-up/down windows, and budget for cold starts on dedicated inference.

    data informada pela fonte; encontrada pelo radar em 10 de set. de 2026

  20. Empresa·oficial·

    Together AI announces strategic partnership with Moonshot AI to natively serve Kimi models

    Together AI partners with Moonshot AI to natively serve Kimi models, starting with the 2.8T parameter Kimi K3, with day zero access and post-training.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  21. Outros·oficial·

    Configuring Dedicated Model Inference

    The three-part resource model behind Together AI Dedicated Model Inference—endpoints, deployments, configs—and how capacity-aware routing ties them together.

    data informada pela fonte; encontrada pelo radar em 10 de set. de 2026

  22. Outros·oficial·

    ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale

    ThunderAgent is a program-aware scheduler for agentic inference. By treating each agent workflow as a schedulable program, it eliminates KV cache thrashing to deliver more than 2x single-node throughput and near-linear multi-node scaling.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  23. Modelos·oficial·

    Kimi K3 vs. GPT-5.6 Sol no DeepSWE: custo, código e roteamento

    Foram executadas 904 rodadas do DeepSWE com Kimi K3 e GPT-5.6 Sol. O Sol lidera no pass@1; o Kimi K3 vence no pass@4 com 2,8× mais soluções por dólar, e o roteamento entre os dois melhora o resultado.

    traduzido por IA · claude-opus-5

    original: “Kimi K3 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing”

    data informada pela fonte; encontrada pelo radar em 10 de set. de 2026

  24. Outros·oficial·

    Kimi K3 vs Claude Fable 5 on DeepSWE: Cost and Coding

    We ran 452 DeepSWE rollouts on Kimi K3 and Claude Fable 5. Fable leads pass@1 by 1.4 points; Kimi K3 wins pass@4 and delivers 2.8x the solves per dollar.

    data informada pela fonte; encontrada pelo radar em 10 de set. de 2026

  25. Outros·oficial·

    The production platform for open-weight AI inference

    Run open models in production with full control over performance, cost, and quality. Deploy in minutes, roll out safely, and scale to your SLOs.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  26. Outros·oficial·

    Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community

    No more two-year compute contracts. Together AI and YC just gave YC startups a faster way to get GPUs.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  27. Outros·oficial·

    What does 99.9% uptime mean for inference?

    Reliability numbers are easy to publish. We break down what 99%, 99.9%, and 99.99% uptime actually require, the failure domains each tier has to survive, and the questions to ask any inference provider before you commit.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  28. Modelos·oficial·

    Together AI recebe o Inkling, novo modelo do Thinking Machines Lab, no dia do lançamento

    A Together AI oferece acesso desde o primeiro dia ao Inkling, modelo multimodal do tipo mistura de especialistas do Thinking Machines Lab, com raciocínio sobre texto, imagem e áudio.

    traduzido por IA · claude-opus-5

    original: “Together AI brings Thinking Machines Lab’s new model Inkling on day 0”

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  29. Outros·oficial·

    New in Together GPU Clusters: Reliability and control for production GPU clusters

    See how Together AI is improving production GPU clusters with passive health checks, node repair, stronger Slurm reliability, OIDC, and startup scripts.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  30. Preços·oficial·

    Provisioned Throughput: capacidade reservada com preço previsível

    O Provisioned Throughput dá capacidade de inferência reservada para modelos abertos de fronteira como MiniMax M3 e GLM-5.2, com preço por token e SLA de 99% de disponibilidade.

    traduzido por IA · claude-opus-5

    original: “Open, convenient and predictable: Introducing Provisioned Throughput”

    data informada pela fonte; encontrada pelo radar em 10 de set. de 2026

  31. Empresa·oficial·

    Announcing our $800M Series C to accelerate the shift to open-source AI

    We raised $800M to accelerate the shift to open-source AI. Here's why the economics of closed models don't scale, and what we're building next.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  32. Pesquisa·oficial·

    Together AI at ICML 2026: frontier research across the full stack

    Nine papers at ICML 2026 across the full stack. The research that becomes the Together platform. Find us at booth B714 in Seoul.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  33. Outros·oficial·

    ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)

    ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves under a third, but a few generated kernels beat any public implementation.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  34. Outros·oficial·

    Kimi K2.7 Code vs Claude Fable 5: Landing pages that cost 94% less

    We generated 12 landing pages with Kimi K2.7 Code and Claude Fable 5. Kimi cost 94% less and scored within a few points on every page. Here's what actually moved the needle.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  35. Outros·oficial·

    Building trust in enterprise AI: Together AI earns ISO 27001:2022 certification

    Together AI has earned ISO 27001:2022 certification, validating our commitment to enterprise-grade security for production AI workloads.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  36. Outros·oficial·

    Serving MiniMax-M3 for efficient inference: Unlocking 1M-Token Context and Multimodality Without Regrets

    How Together served MiniMax-M3 efficiently with KV-block-major sparse attention, paged MSA decode, optimized index scoring, and a Rust-based multimodal gateway.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  37. Outros·oficial·

    How Together AI built the world’s fastest speech-to-text stack

    Together AI built the fastest speech-to-text stack on Artificial Analysis by treating ASR as a full-path systems problem, not just a GPU inference problem.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  38. Outros·oficial·

    Benchmarking inference at scale: coding agents

    Real-world inference benchmarks for coding agents: 31% more TPS than TensorRT-LLM, 2× better TTFT at saturation, and 76% lower cost than Claude Opus 4.6.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  39. APIs·oficial·

    Together AI and Pearl Research Labs Team Up to Reduce the Cost of AI Inference

    Together AI partners with Pearl Research Labs to launch a discounted Pearl-powered inference endpoint for Gemma-4-31B-it-pearl, using Proof of Useful Work to turn AI workloads into crypto emissions.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

  40. Outros·oficial·

    Violin: An open-source video translation skill that breaks language barriers

    Violin is an open-source AI video translation tool that combines speech recognition, LLM translation, and text-to-speech to make video content accessible across languages.

    data informada pela fonte; encontrada pelo radar em 09 de out. de 2026

Fontes monitoradas

O radar lê estes endereços automaticamente. A ordem é a de prioridade: se a primeira falhar, as seguintes continuam entregando.

← Voltar ao radar