OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show | TechCrunch
Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the art.
TechCrunch · Russell Brandom · https://www.facebook.com/techcrunch
Topics: ModelsInfrastructure
Entities: OpenAIModelsInfrastructure