Kog Enables Processing of 3,000 Tokens Per Second with Standard GPUs
French startup Kog has announced a strategy to enhance AI inference speeds using standard data center GPUs. Kog has adopted an approach that involves designing software and model architecture together to maximize the utilization of NVIDIA and AMD GPUs. The company stated that it has integrated runtime, GPU kernels, and model architecture into an optimized pipeline to improve response speeds to requests from AI agents. According to a technology preview of the Kog inference engine released on May 28, it can process 3,000 output tokens per second per request using eight AMD MI300X GPUs, while eight NVIDIA H200 GPUs under the same conditions recorded 2,100 output tokens per second. Kog is primarily targeting AI coding agents and agent-based workflows. The Laneformer 2B model, released on Hugging Face on June 24, has 2.3 billion parameters and demonstrated performance of 45.1% on HumanEval+ and 51.6% on MBPP+. Kog reported in an August LinkedIn post that it achieved 2,857 tokens per second in a live demo. However, there is controversy regarding the interpretation of speed figures, and it has been pointed out that comparisons are difficult due to variations in model size and hardware. Kog's goal is to achieve low latency on existing GPU infrastructure.
-- Price
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

Flare Lowers Inflation to 3% and Implements Revenue Pool Stage

Cracking 1.33 Trillion Daily Tokens: B.AI Powers the “AI Grid” with Full-Stack Infrastructure to Fuel the Agentic Era

Bank of Russia Launches Digital Ruble on September 1, Strategy Buys Bitcoin for $370 Million

PostGREShell: flaw in PostgreSQL turned backup accounts into backdoors

Binance continues EU operations without MiCA license

AEON Launches Agentic Checkout Feature Supported by AI Card

Vice President Vance Asks Central Bank for Interest Rate Cut

Tether's Relentless Pursuit of Justice Takes a Dark Turn

Blackstone Limits Cliffwater Private Loan Fund Redemptions to 5%

European Commission Calls for Agreement on Funding Patriot Missiles for Ukraine

The Cost of AI Tokens in Free Fall: Should We Be Worried?

Tether Blocks 42 Million USDT Before Court Ruling

$317 Billion Stablecoins Become a New Demand Layer for Short-Term Treasuries

Last Night, Silicon Valley Experienced a Battle of AI Titans

Solana inflation cut is premature, SOL Strategies CEO says

Norway Arrests Russian Vessel at Naftogaz's Request Over $4.22 Billion Debt

What the $344M crypto political spending spree wants from Congress next

OpenAI Announces Astra as the First Model with Critical Cybersecurity Capabilities

SEC novel ETF review draws opposition from crypto firms

AI: The Next Billion Crypto Users Will Not Be Human

Prospect Markets, Crypto.com seal deal for U.S. prediction markets platform

Ukraine and Lithuania Strengthen Cooperation to Save Grain Exports

Google Pics Brings AI Images to Workspace, Challenging Canva and Adobe

Flop Labs Launches tclk/1 Protocol for Trustless AI Agent Transactions

Ukraine Needs $27 Billion for Defense

Thai businessmen sue Tether over 42.4M USDT freeze

How recovery of 61 BTC unlocked a potential $432M treasure hunt for early Bitcoin users

Apple Requests Court to Halt OpenAI Hardware Project

Predict Developer Dashboard Launched, Supports Application Creation and Rate Limit Management










