- vendor
- Telnyx
- whatItIs
- API service for running open-weight LLMs (GLM-5.3-Flash, Kimi K3, MiniMax-M3, alongside a still-referenced GLM-5.2) on Telnyx-owned, globally distributed GPUs with OpenAI-compatible endpoints, positioned for low-latency voice AI workloads.
- hosting
- managed-only
- pricingModel
- per-token, starting at $0.21/1M tokens, no GPU rental fees or compute surcharges (per pricing page)
- modelProviders
- open-weight (GLM-5.3-Flash, GLM-5.2, Kimi K3, MiniMax-M3)
- stateModel
- stateless
- toolModel
- OpenAI-compatible API
- includesBrowser
- false
- includesMemory
- false
- longRunning
- per-request
- hitlSupport
- false
- notes
- Telnyx owns the underlying GPU infrastructure rather than renting from cloud providers, with in-region deployment across the Americas, Europe, MENA and APAC (LATAM coming soon, per the product page). Kimi K3 support (1M-token context, native multimodal, prompt caching by default) moved onto Telnyx's own GPUs on 2026-07-28 per Telnyx's release notes (https://telnyx.com/release-notes/kimi-k3-telnyx-inference), migrating from a third-party route. Distinct from tracked open-weight inference vendors (Together AI, Fireworks, Groq, Cerebras, SambaNova, Featherless, SiliconFlow): Telnyx is a comms/API infrastructure company entering inference from an owned-network, voice-latency angle rather than a pure-play inference lab. Correction: the product's hero subheading now reads 'Access GLM-5.3-Flash, Kimi K3, and MiniMax-M3 on dedicated, globally deployed GPUs,' and Telnyx's own GLM-5.3-Flash release note describes it as delivering 'stronger intelligence than GLM-5.2 at a fraction of the cost.' The page is not fully consistent about which GLM version is current: its Agent Runtime section still lists 'GLM-5.2 for dev work' even as the hero promotes GLM-5.3-Flash, so this entry lists both models rather than treating GLM-5.3-Flash as a clean replacement. Maturity kept as ga: the product page itself carries no explicit status label (GA, beta, preview), but Telnyx's own release notes describe each new model (Kimi K3, GLM-5.3-Flash) as now available with live production pricing, so notes no longer contradicts the stored maturity value as it previously did.
- observability
- none
- maturity
- ga
- launched
- 2023-09-10