Together AI

Inferência e fine-tuning de modelos abertos, com anúncios frequentes de preço.

Together API

1 fonte oficial monitorada · última atualização em 21/08/2026

Novidades recentes

40 atualizações

  1. Modelos·oficial·

    GLM-5.3 vs. GPT-5.6 Sol no DeepSWE: custo, código e roteamento

    Foram executadas 904 rodadas do DeepSWE com GLM-5.3 e GPT-5.6 Sol. O Sol lidera no pass@1 por 3,7 pontos; o GLM-5.3 vence no pass@4 pela metade do custo, e uma cascata começando pelo GLM chega a 85,9%.

    traduzido por IA · claude-opus-5

    original: “GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

  2. Features·oficial·

    GLM-5.3 vs. Claude Fable 5 no DeepSWE: custo, código e roteamento

    Foram executadas 904 rodadas do DeepSWE com GLM-5.3 e Claude Fable 5. Empate no pass@1, mas o GLM-5.3 vence no pass@4 e custa 5,4× menos: US$ 3,99 por rodada contra US$ 21,63.

    traduzido por IA · claude-opus-5

    original: “GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

  3. Modelos·oficial·

    DeepSeek V4 Pro 0813 vs. GPT-5.6 Sol no DeepSWE: custo, código e roteamento

    Foram executadas 904 rodadas do DeepSWE com DeepSeek V4 Pro 0813 e GPT-5.6 Sol. O Sol lidera no pass@1 por 10 pontos, mas custa 35× mais; o Pro vence no pass@4, e uma cascata começando pelo Pro chega a 83,0%.

    traduzido por IA · claude-opus-5

    original: “DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

  4. Outros·oficial·

    DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

    We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.

  5. APIs·oficial·

    Teste A/B de modelos em produção

    Tráfego sombra prova que um candidato é operacionalmente sólido, mas não diz se os usuários gostam mais dele. Faça a divisão no endpoint, e não dentro do seu código.

    traduzido por IA · claude-opus-5

    original: “A/B test models in production

  6. Modelos·oficial·

    DeepSeek-V4 Flash 0731 vs. GPT-5.6 Luna no DeepSWE: custo e código

    Foram executadas 900 rodadas do DeepSWE com DeepSeek-V4 Flash e GPT-5.6 Luna. O Luna lidera no pass@1 por 14 pontos; o DeepSeek entrega 4,8× mais soluções por dólar.

    traduzido por IA · claude-opus-5

    original: “DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding

  7. APIs·oficial·

    Kimi K3: the complete developer guide

    Kimi K3 is the first open 3T-class model. See how it benchmarks, what it costs, and how to call it on the Together AI API, with copy-paste code examples.

  8. Outros·oficial·

    Autoscaling endpoints for LLM inference

    GPU utilization can read healthy while your queue backs up, and a new replica takes minutes to warm. Here's how to pick autoscaling metrics, tune scale-up/down windows, and budget for cold starts on dedicated inference.

  9. Empresa·oficial·

    Together AI announces strategic partnership with Moonshot AI to natively serve Kimi models

    Together AI partners with Moonshot AI to natively serve Kimi models, starting with the 2.8T parameter Kimi K3, with day zero access and post-training.

  10. Outros·oficial·

    Configuring Dedicated Model Inference

    The three-part resource model behind Together AI Dedicated Model Inference—endpoints, deployments, configs—and how capacity-aware routing ties them together.

  11. Outros·oficial·

    ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale

    ThunderAgent is a program-aware scheduler for agentic inference. By treating each agent workflow as a schedulable program, it eliminates KV cache thrashing to deliver more than 2x single-node throughput and near-linear multi-node scaling.

  12. Modelos·oficial·

    Kimi K3 vs. GPT-5.6 Sol no DeepSWE: custo, código e roteamento

    Foram executadas 904 rodadas do DeepSWE com Kimi K3 e GPT-5.6 Sol. O Sol lidera no pass@1; o Kimi K3 vence no pass@4 com 2,8× mais soluções por dólar, e o roteamento entre os dois melhora o resultado.

    traduzido por IA · claude-opus-5

    original: “Kimi K3 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  13. Outros·oficial·

    Kimi K3 vs Claude Fable 5 on DeepSWE: Cost and Coding

    We ran 452 DeepSWE rollouts on Kimi K3 and Claude Fable 5. Fable leads pass@1 by 1.4 points; Kimi K3 wins pass@4 and delivers 2.8x the solves per dollar.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  14. Outros·oficial·

    The production platform for open-weight AI inference

    Run open models in production with full control over performance, cost, and quality. Deploy in minutes, roll out safely, and scale to your SLOs.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  15. Outros·oficial·

    Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community

    No more two-year compute contracts. Together AI and YC just gave YC startups a faster way to get GPUs.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  16. Outros·oficial·

    What does 99.9% uptime mean for inference?

    Reliability numbers are easy to publish. We break down what 99%, 99.9%, and 99.99% uptime actually require, the failure domains each tier has to survive, and the questions to ask any inference provider before you commit.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  17. Modelos·oficial·

    Together AI recebe o Inkling, novo modelo do Thinking Machines Lab, no dia do lançamento

    A Together AI oferece acesso desde o primeiro dia ao Inkling, modelo multimodal do tipo mistura de especialistas do Thinking Machines Lab, com raciocínio sobre texto, imagem e áudio.

    traduzido por IA · claude-opus-5

    original: “Together AI brings Thinking Machines Lab’s new model Inkling on day 0

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  18. Outros·oficial·

    New in Together GPU Clusters: Reliability and control for production GPU clusters

    See how Together AI is improving production GPU clusters with passive health checks, node repair, stronger Slurm reliability, OIDC, and startup scripts.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  19. Preços·oficial·

    Provisioned Throughput: capacidade reservada com preço previsível

    O Provisioned Throughput dá capacidade de inferência reservada para modelos abertos de fronteira como MiniMax M3 e GLM-5.2, com preço por token e SLA de 99% de disponibilidade.

    traduzido por IA · claude-opus-5

    original: “Open, convenient and predictable: Introducing Provisioned Throughput

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  20. Empresa·oficial·

    Announcing our $800M Series C to accelerate the shift to open-source AI

    We raised $800M to accelerate the shift to open-source AI. Here's why the economics of closed models don't scale, and what we're building next.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  21. Pesquisa·oficial·

    Together AI at ICML 2026: frontier research across the full stack

    Nine papers at ICML 2026 across the full stack. The research that becomes the Together platform. Find us at booth B714 in Seoul.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  22. Outros·oficial·

    ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)

    ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves under a third, but a few generated kernels beat any public implementation.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  23. Outros·oficial·

    Kimi K2.7 Code vs Claude Fable 5: Landing pages that cost 94% less

    We generated 12 landing pages with Kimi K2.7 Code and Claude Fable 5. Kimi cost 94% less and scored within a few points on every page. Here's what actually moved the needle.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  24. Outros·oficial·

    Building trust in enterprise AI: Together AI earns ISO 27001:2022 certification

    Together AI has earned ISO 27001:2022 certification, validating our commitment to enterprise-grade security for production AI workloads.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  25. Outros·oficial·

    Serving MiniMax-M3 for efficient inference: Unlocking 1M-Token Context and Multimodality Without Regrets

    How Together served MiniMax-M3 efficiently with KV-block-major sparse attention, paged MSA decode, optimized index scoring, and a Rust-based multimodal gateway.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  26. Outros·oficial·

    How Together AI built the world’s fastest speech-to-text stack

    Together AI built the fastest speech-to-text stack on Artificial Analysis by treating ASR as a full-path systems problem, not just a GPU inference problem.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  27. Outros·oficial·

    Benchmarking inference at scale: coding agents

    Real-world inference benchmarks for coding agents: 31% more TPS than TensorRT-LLM, 2× better TTFT at saturation, and 76% lower cost than Claude Opus 4.6.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  28. APIs·oficial·

    Together AI and Pearl Research Labs Team Up to Reduce the Cost of AI Inference

    Together AI partners with Pearl Research Labs to launch a discounted Pearl-powered inference endpoint for Gemma-4-31B-it-pearl, using Proof of Useful Work to turn AI workloads into crypto emissions.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  29. Outros·oficial·

    Violin: An open-source video translation skill that breaks language barriers

    Violin is an open-source AI video translation tool that combines speech recognition, LLM translation, and text-to-speech to make video content accessible across languages.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  30. Produtos·oficial·

    Introducing voice finder — a new tool to quickly find the right voice for your app from over 600+ voices

    Voice finder helps developers search, match, filter, and audition 600+ voices across Together AI TTS models using natural-language prompts or uploaded audio samples.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  31. APIs·oficial·

    Serving DeepSeek-V4: why million-token context is an inference systems problem

    DeepSeek-V4 makes million-token context a serving-systems problem. Together AI explores the inference work behind V4 on NVIDIA HGX B200, including compressed KV layouts, prefix caching, kernel maturity, and endpoint profiles for long-context workloads.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  32. Releases·oficial·

    Deploy and inference any model from HuggingFace

    Learn how to deploy any Hugging Face model in one session using Goose and Together's Dedicated Container Inference. Skip the setup complexity — one prompt gets your model running in a production-grade GPU environment on release day.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  33. Pesquisa·oficial·

    Foundational research powering efficient inference at scale

    As AI moves from research to production, the challenge for AI-native teams shifts from building models to running them — efficiently, reliably, and at scale.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  34. Outros·oficial·

    From 732 bytes to nowhere: shutting down Copy Fail in production

    How Together AI responded to the Copy Fail Linux kernel bug (CVE-2026-31431): disabling the affected crypto interface fleet wide and patching safely.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  35. Empresa·oficial·

    Announcing Together AI and Adaption Partnership

    Together AI and Adaption partner to bring Together Fine-Tuning natively into Adaptive Data, helping teams optimize datasets, run fine-tuning, evaluate results, and deploy stronger open models.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  36. Preços·oficial·

    DeepSeek-V4 Pro chega à Together AI

    O DeepSeek-V4 Pro já está disponível na Together AI com contexto de 512K, modos de raciocínio controláveis e preço com cache de entrada para raciocínio sobre contexto longo.

    traduzido por IA · claude-opus-5

    original: “DeepSeek-V4 Pro now available on Together AI

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  37. Outros·oficial·

    Together AI Brings NVIDIA Nemotron 3 Nano Omni to Developers on Day 0

    NVIDIA Nemotron 3 Nano Omni is now on Together AI: a single open model that reasons across video, images, audio, and text, built for agentic workloads at scale.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  38. Features·oficial·

    Accelerate RL rollouts by up to 50% with distribution-aware speculative decoding

    Rollout is the silent bottleneck in RL post-training. DAS fixes it with adaptive speculative decoding — up to 50% faster, zero degradation in reward quality.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  39. Outros·oficial·

    Capacity without conflict: A guide to multi-tenant GPU cluster design for AI-native teams

    Learn how AI-native companies design multi-tenant GPU clusters that pool capacity without sacrificing team isolation — and how Together AI makes it work in practice.

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

  40. Outros·oficial·

    Parcae: Doing more with fewer parameters using stable looped models

    Parcae is a stable looped language model that matches the quality of a Transformer twice its size — a 770M model reaching 1.3B-level performance. We introduce the first scaling laws for looping and show that increasing recurrence, not just data, is a compute-efficient path to bet

    data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026

Fontes monitoradas

O radar lê estes endereços automaticamente. A ordem é a de prioridade: se a primeira falhar, as seguintes continuam entregando.

← Voltar ao radar