Together AI
Inferência e fine-tuning de modelos abertos, com anúncios frequentes de preço.
Together API
1 fonte oficial monitorada · última atualização em 21/08/2026
Novidades recentes
40 atualizações
- Modelos·oficial·
GLM-5.3 vs. GPT-5.6 Sol no DeepSWE: custo, código e roteamento
Foram executadas 904 rodadas do DeepSWE com GLM-5.3 e GPT-5.6 Sol. O Sol lidera no pass@1 por 3,7 pontos; o GLM-5.3 vence no pass@4 pela metade do custo, e uma cascata começando pelo GLM chega a 85,9%.
traduzido por IA · claude-opus-5
original: “GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing”
- Features·oficial·
GLM-5.3 vs. Claude Fable 5 no DeepSWE: custo, código e roteamento
Foram executadas 904 rodadas do DeepSWE com GLM-5.3 e Claude Fable 5. Empate no pass@1, mas o GLM-5.3 vence no pass@4 e custa 5,4× menos: US$ 3,99 por rodada contra US$ 21,63.
traduzido por IA · claude-opus-5
original: “GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing”
- Modelos·oficial·
DeepSeek V4 Pro 0813 vs. GPT-5.6 Sol no DeepSWE: custo, código e roteamento
Foram executadas 904 rodadas do DeepSWE com DeepSeek V4 Pro 0813 e GPT-5.6 Sol. O Sol lidera no pass@1 por 10 pontos, mas custa 35× mais; o Pro vence no pass@4, e uma cascata começando pelo Pro chega a 83,0%.
traduzido por IA · claude-opus-5
original: “DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing”
- Outros·oficial·
DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.
- APIs·oficial·
Teste A/B de modelos em produção
Tráfego sombra prova que um candidato é operacionalmente sólido, mas não diz se os usuários gostam mais dele. Faça a divisão no endpoint, e não dentro do seu código.
traduzido por IA · claude-opus-5
original: “A/B test models in production”
- Modelos·oficial·
DeepSeek-V4 Flash 0731 vs. GPT-5.6 Luna no DeepSWE: custo e código
Foram executadas 900 rodadas do DeepSWE com DeepSeek-V4 Flash e GPT-5.6 Luna. O Luna lidera no pass@1 por 14 pontos; o DeepSeek entrega 4,8× mais soluções por dólar.
traduzido por IA · claude-opus-5
original: “DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding”
- APIs·oficial·
Kimi K3: the complete developer guide
Kimi K3 is the first open 3T-class model. See how it benchmarks, what it costs, and how to call it on the Together AI API, with copy-paste code examples.
- Outros·oficial·
Autoscaling endpoints for LLM inference
GPU utilization can read healthy while your queue backs up, and a new replica takes minutes to warm. Here's how to pick autoscaling metrics, tune scale-up/down windows, and budget for cold starts on dedicated inference.
- Empresa·oficial·
Together AI announces strategic partnership with Moonshot AI to natively serve Kimi models
Together AI partners with Moonshot AI to natively serve Kimi models, starting with the 2.8T parameter Kimi K3, with day zero access and post-training.
- Outros·oficial·
Configuring Dedicated Model Inference
The three-part resource model behind Together AI Dedicated Model Inference—endpoints, deployments, configs—and how capacity-aware routing ties them together.
- Outros·oficial·
ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale
ThunderAgent is a program-aware scheduler for agentic inference. By treating each agent workflow as a schedulable program, it eliminates KV cache thrashing to deliver more than 2x single-node throughput and near-linear multi-node scaling.
- Modelos·oficial·
Kimi K3 vs. GPT-5.6 Sol no DeepSWE: custo, código e roteamento
Foram executadas 904 rodadas do DeepSWE com Kimi K3 e GPT-5.6 Sol. O Sol lidera no pass@1; o Kimi K3 vence no pass@4 com 2,8× mais soluções por dólar, e o roteamento entre os dois melhora o resultado.
traduzido por IA · claude-opus-5
original: “Kimi K3 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing”
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Outros·oficial·
Kimi K3 vs Claude Fable 5 on DeepSWE: Cost and Coding
We ran 452 DeepSWE rollouts on Kimi K3 and Claude Fable 5. Fable leads pass@1 by 1.4 points; Kimi K3 wins pass@4 and delivers 2.8x the solves per dollar.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Outros·oficial·
The production platform for open-weight AI inference
Run open models in production with full control over performance, cost, and quality. Deploy in minutes, roll out safely, and scale to your SLOs.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Outros·oficial·
Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community
No more two-year compute contracts. Together AI and YC just gave YC startups a faster way to get GPUs.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Outros·oficial·
What does 99.9% uptime mean for inference?
Reliability numbers are easy to publish. We break down what 99%, 99.9%, and 99.99% uptime actually require, the failure domains each tier has to survive, and the questions to ask any inference provider before you commit.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Modelos·oficial·
Together AI recebe o Inkling, novo modelo do Thinking Machines Lab, no dia do lançamento
A Together AI oferece acesso desde o primeiro dia ao Inkling, modelo multimodal do tipo mistura de especialistas do Thinking Machines Lab, com raciocínio sobre texto, imagem e áudio.
traduzido por IA · claude-opus-5
original: “Together AI brings Thinking Machines Lab’s new model Inkling on day 0”
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Outros·oficial·
New in Together GPU Clusters: Reliability and control for production GPU clusters
See how Together AI is improving production GPU clusters with passive health checks, node repair, stronger Slurm reliability, OIDC, and startup scripts.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Preços·oficial·
Provisioned Throughput: capacidade reservada com preço previsível
O Provisioned Throughput dá capacidade de inferência reservada para modelos abertos de fronteira como MiniMax M3 e GLM-5.2, com preço por token e SLA de 99% de disponibilidade.
traduzido por IA · claude-opus-5
original: “Open, convenient and predictable: Introducing Provisioned Throughput”
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Empresa·oficial·
Announcing our $800M Series C to accelerate the shift to open-source AI
We raised $800M to accelerate the shift to open-source AI. Here's why the economics of closed models don't scale, and what we're building next.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Pesquisa·oficial·
Together AI at ICML 2026: frontier research across the full stack
Nine papers at ICML 2026 across the full stack. The research that becomes the Together platform. Find us at booth B714 in Seoul.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Outros·oficial·
ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)
ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves under a third, but a few generated kernels beat any public implementation.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Outros·oficial·
Kimi K2.7 Code vs Claude Fable 5: Landing pages that cost 94% less
We generated 12 landing pages with Kimi K2.7 Code and Claude Fable 5. Kimi cost 94% less and scored within a few points on every page. Here's what actually moved the needle.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Outros·oficial·
Building trust in enterprise AI: Together AI earns ISO 27001:2022 certification
Together AI has earned ISO 27001:2022 certification, validating our commitment to enterprise-grade security for production AI workloads.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Outros·oficial·
Serving MiniMax-M3 for efficient inference: Unlocking 1M-Token Context and Multimodality Without Regrets
How Together served MiniMax-M3 efficiently with KV-block-major sparse attention, paged MSA decode, optimized index scoring, and a Rust-based multimodal gateway.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Outros·oficial·
How Together AI built the world’s fastest speech-to-text stack
Together AI built the fastest speech-to-text stack on Artificial Analysis by treating ASR as a full-path systems problem, not just a GPU inference problem.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Outros·oficial·
Benchmarking inference at scale: coding agents
Real-world inference benchmarks for coding agents: 31% more TPS than TensorRT-LLM, 2× better TTFT at saturation, and 76% lower cost than Claude Opus 4.6.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- APIs·oficial·
Together AI and Pearl Research Labs Team Up to Reduce the Cost of AI Inference
Together AI partners with Pearl Research Labs to launch a discounted Pearl-powered inference endpoint for Gemma-4-31B-it-pearl, using Proof of Useful Work to turn AI workloads into crypto emissions.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Outros·oficial·
Violin: An open-source video translation skill that breaks language barriers
Violin is an open-source AI video translation tool that combines speech recognition, LLM translation, and text-to-speech to make video content accessible across languages.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Produtos·oficial·
Introducing voice finder — a new tool to quickly find the right voice for your app from over 600+ voices
Voice finder helps developers search, match, filter, and audition 600+ voices across Together AI TTS models using natural-language prompts or uploaded audio samples.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- APIs·oficial·
Serving DeepSeek-V4: why million-token context is an inference systems problem
DeepSeek-V4 makes million-token context a serving-systems problem. Together AI explores the inference work behind V4 on NVIDIA HGX B200, including compressed KV layouts, prefix caching, kernel maturity, and endpoint profiles for long-context workloads.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Releases·oficial·
Deploy and inference any model from HuggingFace
Learn how to deploy any Hugging Face model in one session using Goose and Together's Dedicated Container Inference. Skip the setup complexity — one prompt gets your model running in a production-grade GPU environment on release day.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Pesquisa·oficial·
Foundational research powering efficient inference at scale
As AI moves from research to production, the challenge for AI-native teams shifts from building models to running them — efficiently, reliably, and at scale.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Outros·oficial·
From 732 bytes to nowhere: shutting down Copy Fail in production
How Together AI responded to the Copy Fail Linux kernel bug (CVE-2026-31431): disabling the affected crypto interface fleet wide and patching safely.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Empresa·oficial·
Announcing Together AI and Adaption Partnership
Together AI and Adaption partner to bring Together Fine-Tuning natively into Adaptive Data, helping teams optimize datasets, run fine-tuning, evaluate results, and deploy stronger open models.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Preços·oficial·
DeepSeek-V4 Pro chega à Together AI
O DeepSeek-V4 Pro já está disponível na Together AI com contexto de 512K, modos de raciocínio controláveis e preço com cache de entrada para raciocínio sobre contexto longo.
traduzido por IA · claude-opus-5
original: “DeepSeek-V4 Pro now available on Together AI”
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Outros·oficial·
Together AI Brings NVIDIA Nemotron 3 Nano Omni to Developers on Day 0
NVIDIA Nemotron 3 Nano Omni is now on Together AI: a single open model that reasons across video, images, audio, and text, built for agentic workloads at scale.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Features·oficial·
Accelerate RL rollouts by up to 50% with distribution-aware speculative decoding
Rollout is the silent bottleneck in RL post-training. DAS fixes it with adaptive speculative decoding — up to 50% faster, zero degradation in reward quality.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Outros·oficial·
Capacity without conflict: A guide to multi-tenant GPU cluster design for AI-native teams
Learn how AI-native companies design multi-tenant GPU clusters that pool capacity without sacrificing team isolation — and how Together AI makes it work in practice.
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
- Outros·oficial·
Parcae: Doing more with fewer parameters using stable looped models
Parcae is a stable looped language model that matches the quality of a Transformer twice its size — a 770M model reaching 1.3B-level performance. We introduce the first scaling laws for looping and show that increasing recurrence, not just data, is a compute-efficient path to bet
data informada pela fonte; encontrada pelo radar em 25 de ago. de 2026
Fontes monitoradas
O radar lê estes endereços automaticamente. A ordem é a de prioridade: se a primeira falhar, as seguintes continuam entregando.
- 1. https://www.together.ai/blog/rss.xml (rss)