Together AI does not target individual consumer chatbots or casual users; it is a cloud platform built specifically for AI researchers, developers, and enterprise teams.
Key Details
• Target Audience: AI engineers, software developers, researchers, and companies building custom applications.
• Core Offerings: Cloud computing power (GPUs), serverless and dedicated model inference, fine-tuning, and pre-training infrastructure for open-source AI models.
• Pricing & Access: It is not a free or consumer-facing subscription service; access requires a prepaid account with a minimum $5 credit purchase via Together AI Docs.
• Missing Feature: It does not offer a ready-to-use, general-purpose consumer assistant app (like ChatGPT or Claude).
"Paris, early July 2025. At the Raise Summit industry gathering, Vipul Ved Prakash—founder and CEO of the AI provider Together AI—demonstrates to the audience exactly what his engineers have achieved with the Chinese language model DeepSeek. Initially, processing a million tokens through the model costs eight dollars. Then, his team tweaks the software—optimizing it compute core by compute core—until the cost drops to 55 cents. The next generation of chips, Prakash notes, will slash costs by another factor of three to four.
A year later, that calculation on stage translates into a valuation. On July 1, Together AI announced an $800 million funding round that values the San Francisco-based company at $8.3 billion—more than double its valuation from February 2025. The round is led by Aramco Ventures, the investment arm of the Saudi state-owned giant. Nvidia is investing again, joined by backers including Vista Equity Partners, General Catalyst, Emergence Capital, Salesforce Ventures, and the Taiwanese contract manufacturer Pegatron.
At first glance, the underlying business model seems mundane. Together does not build its own cutting-edge models; instead, it operates models created by others—such as DeepSeek, Kimi, and MiniMax from China, as well as Nvidia’s Nemotron and Meta’s Llama. The company runs these models on tens of thousands of Nvidia chips and sells the active processing—known in the industry as "inference"—at the lowest possible price.
Intelligence is becoming a fundamental economic resource, Prakash wrote regarding the funding round—as indispensable as electricity, bandwidth, or capital. "Our mission is to ensure that intelligence is abundant, not expensive." "The future of AI won't belong to just a handful of companies; it is being built by millions of developers and businesses."
Prakash has previously rattled an industry with software he gave away for free. In the late 1990s, the St. Stephen's College (Delhi) dropout wrote a spam filter called Vipul's Razor; it aggregated user reports into "fingerprints" and redistributed them to everyone. The program ended up on more than ten million servers, and in 2003, *MIT Technology Review* named its creator one of the top innovators under 35. The filter was commercialized by Cloudmark, a company Prakash founded with Jordan Ritter—an engineer who had previously helped build Napster. Next came Topsy Labs, a search engine for the Twitter data stream, which Apple acquired in 2013 for a reported $200 million. Prakash subsequently spent five years at Apple leading a team of over 250 engineers building search technology for Spotlight, Safari, and Siri.
In June 2022, he founded Together alongside systems researcher Ce Zhang and Stanford professors Chris Ré and Percy Liang. The company now presents a broader version of its founding story: its website lists five founders, including Princeton professor Tri Dao—inventor of the Flash Attention computational method, which is used by virtually every major AI lab. However, Dao did not join as Chief Scientist until July 2023. In the initial blog post announcing the launch, Prakash had written simply of "Chris, Percy, Ce, and myself." It was a minor revision with a clear message: what is being sold here is research, not merely computing power.
The actual price difference is evident in the rate cards. In early July,
Together AI charged $1.74 per million input tokens and $3.48 for output tokens for its top-tier open model, Deepseek V4 Pro.
Anthropic, meanwhile, set the price for Claude Opus $25 for GPT-4.5 and $35 for OpenAI’s GPT-5.5.
Depending on the usage mix, this results in a factor of six to nine—for a model that ranks one tier below the market leaders in independent coding tests, but not two.
The "up to 60-fold" cost reduction promised in company marketing is only realized by those who switch to the low-cost offerings of Chinese providers or opt for stripped-down model variants. The actual savings are in the single-digit range. That is enough.
Customers make similar calculations. Customer service specialist Decagon reports that its costs per conversation turn dropped to one-sixth after switching to Together, compared to OpenAI’s smaller GPT-5 mini model.
Vineet Khosla, CTO of the *Washington Post*, cites full performance at lower costs than those of closed-ecosystem providers—while meeting strict data privacy requirements—in Together’s customer testimonials.
Caiming Xiong, Head of Research at Salesforce, reports response times cut in half alongside savings of around one-third. And Cursor, arguably the industry’s best-known coding assistant, runs part of its workload on 72 Blackwell-generation chips at Together.
The figure Together uses in its marketing is $1.15 billion: The level of annual bookings recorded in the last quarter. Bookings are contracts signed, but no revenue actually collected. The analytics firm Sacra estimates actual annualized revenue at around one billion dollars, with an estimated gross margin of 45 percent. Old-school software companies operate with gross margins of 80 to 90 percent. By this assessment, the token sale itself runs roughly at cost; the real money is made by renting out entire chip clusters. If you want to make money using someone else's free software, you end up with Red Hat, not Microsoft.
Within its peer group, the valuation seems almost conservative. Rival Baseten raised $1.5 billion in June at a $13 billion valuation; Fireworks is reportedly negotiating a deal at $15 billion; and chipmaker Cerebras launched an IPO in May valuing the company at more than fifty times its revenue. Compared to such multiples, the figure of roughly eight times annual revenue derived from Sacra’s estimate seems almost modest. Yet a warning sign lies hidden in the fine print: in March, the company reported 27 customer contracts worth over one million dollars each, and one worth over a billion. It does not disclose how much of the total booked value is tied to that single contract.
The identity of the lead investor reveals more about the state of the AI economy than any valuation figure. Prosperity 7 is the program within Aramco Ventures that is spearheading the investment in Together. The name alludes to the specific well that launched the Saudi oil era in 1938. Abhishek Shukla, the program's US head, calls the development of AI infrastructure "the greatest infrastructure project in human history." The fund is well acquainted with American sensitivities from firsthand experience: in 2023, the Committee on Foreign Investment in the United States (CFIUS) forced Prosperity 7 to sell its stake in the chip startup Rain AI. No such review regarding the investment in Together has been reported to date; Together sells software, not chips.
The second twist lies in the underlying mechanics. A large proportion of the open models Together promotes originate in China: Deepseek, Moonshot’s Kimi, Minimax, and Alibaba’s Qwen. On the developer platform Open Router—whose data Together itself cites—usage of open models tripled within a single year. At times, Deepseek alone accounted for roughly one-sixth of all traffic there. The model weights run on American servers. Requests do not need to leave the country—unlike with the Deepseek app, where data flows to China. Washington has not taken issue with this so far. It has, however, objected to access regarding its own labs: in June, the US Department of Commerce ordered that foreigners be blocked from accessing Anthropic’s cutting-edge models—marking the first instance of export controls being applied to an AI model rather than chips. The agency largely rescinded the order after 18 days. The episode did not affect the Chinese models hosted on Together’s American servers.
However, anyone who believes price dictates the market underestimates major corporate clients. In late 2025, the venture capital firm Menlo Ventures released its annual survey of large enterprises. Open models accounted for 11 percent of AI usage in the survey, down from 19 percent the previous year; Anthropic, OpenAI, and Google shared 88 percent of the market between them. Chinese open models—Together’s showcase products—stood at just one percent among these corporations. The research institute Epoch AI, meanwhile, recently measured a lag of around four months between open models and the closed-source leaders—a gap that appears to be widening. And for long, chained agent tasks—currently the most expensive use case—programmers predominantly turn to Anthropic’s Claude anyway. Together is thus growing in the cost-conscious segment of the market, while the leading labs defend the rest.
Nebius—which spun out of the Yandex group and is now based in Amsterdam—is Together’s European counterpart; it is listed on the Nasdaq. The company rents out the same Nvidia chips and runs the same open models, but markets itself based on European legal compliance rather than price. Nvidia is investing here as well, having put in two billion dollars in March. Additionally, Germany’s Ionos and France’s OVH Cloud and Scaleway host open models on European servers for customers who prioritize jurisdiction over saving the last cent per token.
The fresh capital is flowing primarily into infrastructure—specifically, facilities and chips. Around 200 megawatts of grid capacity have been documented so far, with sites in Maryland, Sweden, and—coming soon—Memphis. Since the funding round, the company has spoken of commitments totaling over 500 megawatts, which its backers are expected to finance independently of Together’s own balance sheet, according to the company, capacity is set to grow roughly fifty-fold within five years.
No company with an eight-billion-dollar valuation could shoulder that alone; expansion is being driven by partners who build and own the facilities, while Together leases and operates them.
Meanwhile, Nvidia is investing in a customer that buys chips in order to rent them out. Prakash does not dispute this circularity. He told the industry news service Axios Pro in early July that he would not rule out an IPO in 2027.
Intelligence is like electricity—that is his favorite analogy. Electrification, of course, produced two types of winners. Electricity became cheap. The ones who got rich were those who owned the grids.” [1]
1. Das Start-up, das Anthropic und Open AI angreift: Together AI verkauft Rechenleistung für fremde Gratismodelle und ist bereits mehr als acht Milliarden Dollar wert. Der größte Geldgeber ist der saudische Ölkonzern Aramco. Frankfurter Allgemeine Zeitung; Frankfurt. 14 July 2026: 19. Von Marcus Schuler, San Francisco
Komentarų nėra:
Rašyti komentarą