Together AI
Inferência e fine-tuning de modelos abertos, com anúncios frequentes de preço.
Together API
1 fonte oficial monitorada · última atualização em 06/10/2026
Novidades recentes
40 atualizações
- Outros·oficial·
Expanding our enterprise inference capacity with IBM Cloud and NVIDIA
Enterprises can now run open models at production scale on a dedicated B300 inference cluster, built by Together AI, IBM Cloud, and NVIDIA
- Outros·oficial·
Together Link: open models in the harness you already use. Start with one command today.
Together Link brings frontier open models like GLM 5.3 and Kimi K3 into the coding agent your team already uses, cutting model spend by over 50%.
- Modelos·oficial·
How to train your own Jev for $17
We just launched our own Jev-like classifier, together/Tev1-4B-experimental, on top of Qwen3.5 4B on Together’s serverless platform. In this blog post we’ll show you how to fine-tune your own version!
- Outros·oficial·
Canary rollouts: upgrade models in production without downtime
A hard model swap exposes every user at once, and rolling back means cold-starting the old deployment under pressure. Here's how staged traffic ramps, metric gates, and automatic rollback work on dedicated inference.
- Outros·oficial·
How a global fintech scaled coding agent traffic with Dedicated Model Inference
Inside a global bank's shift to self-serve dedicated inference: how Together's DMI gave engineering teams direct control over scaling, models, and testing.
- Outros·oficial·
Migrating from closed to open source models, Together
Moving from closed to open source models can take weeks, not years. A five-stage playbook: discover, evaluate, adapt, decide, and production.
- Preços·oficial·
Together AI expands fine-tuning service with more models, live metrics, and finer controls
Together Fine-Tuning adds the latest open-weight models, live experiment tracking, Expert LoRA, early stopping, tokenized dataset previews, pre-flight validation, and lower training prices on selected models.
- Preços·oficial·
Introducing preemptible compute: the same compute, half the price
Together GPU Clusters now supports preemptible compute: the same GPU capacity at a flat 50% of the on-demand rate, with a five-minute drain window.
- Outros·oficial·
To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!
We ported ThunderKittens to NVIDIA's Vera Rubin NVL72 and rebuilt our NVFP4 GEMM around the new hardware, taking it from 42% of roofline to over 22 PFLOPS — competitive with cuBLAS and CuTe DSL. Here is what changed in the ISA and how we used it.
- Outros·oficial·
The Open Source AI Stack
A deep dive into the open model AI stack — model, inference, gateways and routers, harness, and tools — and how keeping each layer independent lets you swap in a new open model in minutes instead of rebuilding your workflow.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Outros·oficial·
GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing
We ran 900 DeepSWE rollouts on GLM-5.3 and GLM-5.3 Flash. Flash gives up 5.6 points of pass@1 at 17x lower cost, and only 2.6 points at pass@4.
- Modelos·oficial·
GLM-5.3 vs. GPT-5.6 Sol no DeepSWE: custo, código e roteamento
Foram executadas 904 rodadas do DeepSWE com GLM-5.3 e GPT-5.6 Sol. O Sol lidera no pass@1 por 3,7 pontos; o GLM-5.3 vence no pass@4 pela metade do custo, e uma cascata começando pelo GLM chega a 85,9%.
traduzido por IA · claude-opus-5
original: “GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing”
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Features·oficial·
GLM-5.3 vs. Claude Fable 5 no DeepSWE: custo, código e roteamento
Foram executadas 904 rodadas do DeepSWE com GLM-5.3 e Claude Fable 5. Empate no pass@1, mas o GLM-5.3 vence no pass@4 e custa 5,4× menos: US$ 3,99 por rodada contra US$ 21,63.
traduzido por IA · claude-opus-5
original: “GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing”
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Modelos·oficial·
DeepSeek V4 Pro 0813 vs. GPT-5.6 Sol no DeepSWE: custo, código e roteamento
Foram executadas 904 rodadas do DeepSWE com DeepSeek V4 Pro 0813 e GPT-5.6 Sol. O Sol lidera no pass@1 por 10 pontos, mas custa 35× mais; o Pro vence no pass@4, e uma cascata começando pelo Pro chega a 83,0%.
traduzido por IA · claude-opus-5
original: “DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing”
- Outros·oficial·
DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- APIs·oficial·
Teste A/B de modelos em produção
Tráfego sombra prova que um candidato é operacionalmente sólido, mas não diz se os usuários gostam mais dele. Faça a divisão no endpoint, e não dentro do seu código.
traduzido por IA · claude-opus-5
original: “A/B test models in production”
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Modelos·oficial·
DeepSeek-V4 Flash 0731 vs. GPT-5.6 Luna no DeepSWE: custo e código
Foram executadas 900 rodadas do DeepSWE com DeepSeek-V4 Flash e GPT-5.6 Luna. O Luna lidera no pass@1 por 14 pontos; o DeepSeek entrega 4,8× mais soluções por dólar.
traduzido por IA · claude-opus-5
original: “DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding”
data informada pela fonte; encontrada pelo radar em 10 de set. de 2026
- APIs·oficial·
Kimi K3: the complete developer guide
Kimi K3 is the first open 3T-class model. See how it benchmarks, what it costs, and how to call it on the Together AI API, with copy-paste code examples.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Outros·oficial·
Autoscaling endpoints for LLM inference
GPU utilization can read healthy while your queue backs up, and a new replica takes minutes to warm. Here's how to pick autoscaling metrics, tune scale-up/down windows, and budget for cold starts on dedicated inference.
data informada pela fonte; encontrada pelo radar em 10 de set. de 2026
- Empresa·oficial·
Together AI announces strategic partnership with Moonshot AI to natively serve Kimi models
Together AI partners with Moonshot AI to natively serve Kimi models, starting with the 2.8T parameter Kimi K3, with day zero access and post-training.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Outros·oficial·
Configuring Dedicated Model Inference
The three-part resource model behind Together AI Dedicated Model Inference—endpoints, deployments, configs—and how capacity-aware routing ties them together.
data informada pela fonte; encontrada pelo radar em 10 de set. de 2026
- Outros·oficial·
ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale
ThunderAgent is a program-aware scheduler for agentic inference. By treating each agent workflow as a schedulable program, it eliminates KV cache thrashing to deliver more than 2x single-node throughput and near-linear multi-node scaling.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Modelos·oficial·
Kimi K3 vs. GPT-5.6 Sol no DeepSWE: custo, código e roteamento
Foram executadas 904 rodadas do DeepSWE com Kimi K3 e GPT-5.6 Sol. O Sol lidera no pass@1; o Kimi K3 vence no pass@4 com 2,8× mais soluções por dólar, e o roteamento entre os dois melhora o resultado.
traduzido por IA · claude-opus-5
original: “Kimi K3 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing”
data informada pela fonte; encontrada pelo radar em 10 de set. de 2026
- Outros·oficial·
Kimi K3 vs Claude Fable 5 on DeepSWE: Cost and Coding
We ran 452 DeepSWE rollouts on Kimi K3 and Claude Fable 5. Fable leads pass@1 by 1.4 points; Kimi K3 wins pass@4 and delivers 2.8x the solves per dollar.
data informada pela fonte; encontrada pelo radar em 10 de set. de 2026
- Outros·oficial·
The production platform for open-weight AI inference
Run open models in production with full control over performance, cost, and quality. Deploy in minutes, roll out safely, and scale to your SLOs.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Outros·oficial·
Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community
No more two-year compute contracts. Together AI and YC just gave YC startups a faster way to get GPUs.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Outros·oficial·
What does 99.9% uptime mean for inference?
Reliability numbers are easy to publish. We break down what 99%, 99.9%, and 99.99% uptime actually require, the failure domains each tier has to survive, and the questions to ask any inference provider before you commit.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Modelos·oficial·
Together AI recebe o Inkling, novo modelo do Thinking Machines Lab, no dia do lançamento
A Together AI oferece acesso desde o primeiro dia ao Inkling, modelo multimodal do tipo mistura de especialistas do Thinking Machines Lab, com raciocínio sobre texto, imagem e áudio.
traduzido por IA · claude-opus-5
original: “Together AI brings Thinking Machines Lab’s new model Inkling on day 0”
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Outros·oficial·
New in Together GPU Clusters: Reliability and control for production GPU clusters
See how Together AI is improving production GPU clusters with passive health checks, node repair, stronger Slurm reliability, OIDC, and startup scripts.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Preços·oficial·
Provisioned Throughput: capacidade reservada com preço previsível
O Provisioned Throughput dá capacidade de inferência reservada para modelos abertos de fronteira como MiniMax M3 e GLM-5.2, com preço por token e SLA de 99% de disponibilidade.
traduzido por IA · claude-opus-5
original: “Open, convenient and predictable: Introducing Provisioned Throughput”
data informada pela fonte; encontrada pelo radar em 10 de set. de 2026
- Empresa·oficial·
Announcing our $800M Series C to accelerate the shift to open-source AI
We raised $800M to accelerate the shift to open-source AI. Here's why the economics of closed models don't scale, and what we're building next.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Pesquisa·oficial·
Together AI at ICML 2026: frontier research across the full stack
Nine papers at ICML 2026 across the full stack. The research that becomes the Together platform. Find us at booth B714 in Seoul.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Outros·oficial·
ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)
ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves under a third, but a few generated kernels beat any public implementation.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Outros·oficial·
Kimi K2.7 Code vs Claude Fable 5: Landing pages that cost 94% less
We generated 12 landing pages with Kimi K2.7 Code and Claude Fable 5. Kimi cost 94% less and scored within a few points on every page. Here's what actually moved the needle.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Outros·oficial·
Building trust in enterprise AI: Together AI earns ISO 27001:2022 certification
Together AI has earned ISO 27001:2022 certification, validating our commitment to enterprise-grade security for production AI workloads.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Outros·oficial·
Serving MiniMax-M3 for efficient inference: Unlocking 1M-Token Context and Multimodality Without Regrets
How Together served MiniMax-M3 efficiently with KV-block-major sparse attention, paged MSA decode, optimized index scoring, and a Rust-based multimodal gateway.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Outros·oficial·
How Together AI built the world’s fastest speech-to-text stack
Together AI built the fastest speech-to-text stack on Artificial Analysis by treating ASR as a full-path systems problem, not just a GPU inference problem.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Outros·oficial·
Benchmarking inference at scale: coding agents
Real-world inference benchmarks for coding agents: 31% more TPS than TensorRT-LLM, 2× better TTFT at saturation, and 76% lower cost than Claude Opus 4.6.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- APIs·oficial·
Together AI and Pearl Research Labs Team Up to Reduce the Cost of AI Inference
Together AI partners with Pearl Research Labs to launch a discounted Pearl-powered inference endpoint for Gemma-4-31B-it-pearl, using Proof of Useful Work to turn AI workloads into crypto emissions.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
- Outros·oficial·
Violin: An open-source video translation skill that breaks language barriers
Violin is an open-source AI video translation tool that combines speech recognition, LLM translation, and text-to-speech to make video content accessible across languages.
data informada pela fonte; encontrada pelo radar em 09 de out. de 2026
Fontes monitoradas
O radar lê estes endereços automaticamente. A ordem é a de prioridade: se a primeira falhar, as seguintes continuam entregando.
- 1. https://www.together.ai/blog/rss.xml (rss)