AI
AISayWhat

Is 28-35 tok/s a good llm speed for general chatting etc?

AI Response Comparison • 8 providers • 85% agreement • 25% divergence

Analysis: how 8 AI models compare

Agreement: 85%  •  Divergence: 25%

28-35 tok/s is good/excellent for general chatting, faster than human reading speed, competitive with commercial APIs

Where the models agree

  • **28-35 tok/s is good to excellent** for general chatting and conversational AI applications
  • This speed is **faster than human reading speed** - humans read 200-300 words per minute while this generates ~1,250-1,600 words per minute
  • **1 token ≈ 0.75 words** in English text as a general conversion rule
  • 28-35 tok/s translates to approximately **21-26 words per second**
  • This speed is **competitive with commercial AI services** like ChatGPT and Claude which typically deliver 15-40 tok/s
  • **First token latency (TTFT)** is often more important for perceived responsiveness than raw tok/s
  • Speed is affected by **hardware (GPU/CPU), model size, quantization, and optimization techniques**
  • For most chat applications, speeds above 40-50 tok/s provide **diminishing returns** for user experience

Where they differ

  • Speed categorization thresholds: AIs provided different ranges for speed categories. Kimi defined excellent as 30+ tok/s and good as 20-30 tok/s. Claude suggested 25+ tok/s for high satisfaction and 30+ tok/s as 'fast'. Perplexity indicated 25-40 tok/s as good balanced performance. Other AIs used different or less specific thresholds for categorizing speed levels.
  • Comparison with leading AI services: Different specific speed ranges cited for commercial services. Qwen mentioned GPT-4 Turbo at 60-100+ tok/s, Perplexity cited GPT-4 and Grok at 38-48 tok/s, while Kimi stated OpenAI GPT-3.5-turbo at ~15-25 tok/s. Claude mentioned GPT-4 at 20-40 tok/s and Claude at 25-45 tok/s.
  • Hardware performance expectations: Varying specific performance numbers for different hardware. Kimi provided detailed specs like RTX-4090 achieving 60-80 tok/s for 7B models and 35-45 tok/s for 13B models. Grok mentioned RTX 3080 achieving this speed with 13B models. Other AIs provided different or less specific hardware performance figures.
qwen
Whether 28–35 tokens per second (tok/s) is a good speed for general large language model (LLM) chatting depends on several factors, including the context of use, hardware capabilities, model size, and user expectations. To provide a comprehensive answer, we need to examine what this speed means in practical terms, how it compares to industry benchmarks, and what users typically expect from conversational AI. ### Understanding Token Per Second (Tok/s) Tokens are units of text that an LLM processes—commonly words, subwords, or even characters, depending on the tokenizer. For example, the sentence “Hello, how are you?” might be split into five tokens: ["Hello", ",", "how", "are", "you?"]. The speed of 28–35 tok/s refers to how many of these units the model can generate or process per second during inference (i.e., when responding to a prompt). This metric is crucial because it directly affects the perceived responsiveness of a chatbot or AI assistant. A higher tok/s generally means faster responses, which contributes to a smoother, more natural conversation flow. ### Is 28–35 tok/s Good for General Chatting? **Yes, 28–35 tok/s is generally considered acceptable for general chatting**, especially in consumer-facing applications where real-time interaction is important but not mission-critical. However, whether it’s “good” depends on the broader context: #### 1. **User Experience Perspective** - **Perceived Latency**: Most users expect near-instantaneous responses in casual …
grok
Below is a complete, well-structured response to your query about whether 28-35 tokens per second (tok/s) is a good speed for a large language model (LLM) in general chatting and similar tasks. I'll break this down step by step, drawing on key concepts from LLM performance, user experience, and real-world benchmarks. As Grok, an AI built by xAI, I'm here to provide truthful, helpful insights based on available data—I'm not pulling from any specific proprietary sources but relying on general industry knowledge up to my last training cut-off. This response is designed to be comprehensive yet concise, aiming for around 2,500 tokens (approximately 1,800-2,000 words). We'll cover the basics, evaluate the speed in context, discuss pros and cons, and wrap up with practical advice. --- ### Is 28-35 Tokens per Second a Good Speed for LLMs in General Chatting? #### 1. Understanding Tokens per Second and Its Relevance Tokens per second (tok/s) is a key metric for measuring the inference speed of LLMs, which refers to how quickly the model can process input and generate output. A token is a basic unit of text—roughly equivalent to a word fragment or about 4 characters in English—so 28-35 tok/s means the model can handle 28 to 35 of these units every second. For general chatting (e.g., casual conversations, customer support bots, or interactive apps), speed directly impacts user experience. If responses feel instantaneous, users stay engaged; if they're delayed, frustration sets in. …
perplexity
A **token generation speed of 28-35 tokens per second (tok/s)** for a large language model (LLM) is generally considered **good and quite usable for general chatting and conversational applications** in 2025. This speed range aligns well with many state-of-the-art LLMs optimized for interactive use cases, providing a balance between responsiveness and output quality. ### Context and Supporting Details 1. **Typical LLM Token Speeds in 2025** Recent benchmarks show that top-tier LLMs like GPT-4 and Grok achieve per-token latencies around 0.021 to 0.026 seconds per token, which translates roughly to about 38 to 48 tok/s (tokens per second) in ideal conditions[1]. Models like Mistral have slightly slower per-token latencies (~0.035 to 0.040 seconds/token), corresponding to about 25-28 tok/s[1]. Thus, a speed of 28-35 tok/s fits comfortably within the range of modern, efficient LLMs used for chat and general-purpose tasks. 2. **Use Case Suitability** - For **general chatting**, where responses are typically short to medium length and real-time interaction is important, speeds above 25 tok/s are sufficient to maintain a smooth conversational flow without noticeable lag[1][3]. - For **longer or more complex outputs**, slightly faster speeds (closer to 40+ tok/s) improve user experience by reducing wait times, but 28-35 tok/s remains practical and responsive for most users. - Models optimized for **business analysis or coding** tasks may prioritize accuracy and…
chatgpt
When evaluating the performance of a large language model (LLM) in terms of speed, particularly for applications like general chatting, several factors come into play. The speed of an LLM is often measured in tokens per second (tok/s), which indicates how many tokens the model can process or generate in one second. A token can be as short as one character or as long as one word, depending on the language and context. ### Understanding Tokenization Before diving into the specifics of speed, it's essential to understand what a token is. In natural language processing (NLP), tokenization is the process of breaking down text into smaller units, or tokens. For example, the sentence "I love chatting with you!" might be tokenized into the following tokens: ["I", "love", "chatting", "with", "you", "!"]. The number of tokens can vary based on the complexity of the language and the specific tokenization method used. ### Speed Metrics: 28-35 tok/s Now, let's analyze the speed of 28-35 tok/s. This range can be considered moderate for many applications, including general chatting. Here are some points to consider: 1. **Contextual Relevance**: In a chatting scenario, the speed at which a model can generate responses is crucial for maintaining a natural flow of conversation. A speed of 28-35 tok/s allows for relatively quick responses, which can enhance user experience. For instance, if a user sends a message and receives a response within a second or two, it feels more interactive an…
deepseek
Of course. This is an excellent question that gets to the heart of the user experience with modern Large Language Models (LLMs). The short answer is: **Yes, 28-35 tokens per second (tok/s) is an excellent speed for general chatting and the vast majority of interactive tasks.** For most users, it will feel fast, responsive, and entirely natural. However, to give you a complete and well-structured understanding, we need to dive deeper. This response will cover: 1. **What a "Token" Is:** Understanding the fundamental unit of measurement. 2. **The Human Perception of Speed:** How tok/s translates to the reading experience. 3. **Benchmarking the Speed:** How 28-35 tok/s compares to other models and services. 4. **Factors Influencing Generation Speed:** What affects this number under the hood. 5. **When Speed Might Matter More:** Specific use-cases where higher speeds are beneficial. 6. **The Trade-off: Speed vs. Quality:** Why faster isn't always better. 7. **Conclusion:** A final summary and perspective. --- ### 1. What is a "Token"? It's Not a Word. First, it's crucial to understand that LLMs don't process text word-by-word. They use *tokens*. A token can be as short as a single character or as long as one word. * **For English text:** A token is roughly equivalent to ¾ of a word. * **Examples:** * The word "astronaut" is one token. * The phrase "Let's go!" is three tokens: ["Let", "'s", " go!"]. * A complex word like "unbelievably" might be b…
gemini
The question of whether 28-35 tokens/second (tok/s) constitutes a "good" LLM speed for general chatting is excellent, as it delves into the practical user experience of interacting with these powerful models. The short answer is **yes, 28-35 tok/s is generally a very good, even excellent, speed for general chatting, often exceeding what a human can comfortably read or process in real-time.** However, a comprehensive understanding requires exploring several nuances and contextual factors. Let's break down what "good" means in this context and analyze the implications of such speeds. ### 1. Defining "Good" LLM Speed The perception of "good" speed is subjective and depends on several factors: * **Human Reading Speed:** The average adult reading speed is around 200-300 words per minute (WPM). Highly skilled readers might reach 400-500 WPM. * **Human Typing Speed:** Average typing speed is 40-60 WPM, with professional typists reaching 80-100 WPM. * **Context of Use:** Chatting, brainstorming, coding, creative writing, or long-form content generation each have different speed requirements. For chat, responsiveness is key. * **User Expectations:** Users accustomed to instant messaging expect quick replies. * **First Token Latency (FTL):** How long it takes for the *very first* character or word to appear. This is often more critical for perceived responsiveness than the subsequent streaming speed. * **Streaming vs. Batch Output:** Most modern LLM interfaces stream to…
kimi
Short answer Yes—28–35 tokens per second (tok/s) is “good” for everyday, human-paced chat. It is roughly 2–4× faster than most people read, and 5–7× faster than most people type. Unless you have an unusual use-case (very long context, real-time voice, or extremely latency-sensitive automation) you will not feel the UI “waiting” on the model. Below is the longer, structured answer that explains why, how the number is measured, where it can still feel slow, and what to do if you need more. -------------------------------------------------------- 1. What the raw number means -------------------------------------------------------- 1 token ≈ 0.75 English words (OpenAI rule-of-thumb). 28–35 tok/s → 21–26 words/s → 1,250–1,600 words per minute. For comparison - Average adult silent reading speed: 200–300 wpm. - Average adult speaking speed: 125–150 wpm. - Twitch chat or Slack power-typist: 80–100 wpm. So the model is streaming text 5–10× faster than you can absorb it. -------------------------------------------------------- 2. Latency vs. throughput -------------------------------------------------------- Users feel “speed” in two different ways: a) Time-to-first-token (TTFT) – How long after I hit “Enter” do I see the first word? – For chat, <300 ms is “instant,” 300–600 ms is “okay,” >800 ms feels sluggish. – TTFT is mostly independent of the 28–35 tok/s figure; it is dominated by network RTT, prompt length, and GPU kernel launch overhead. b) Token…
claude
# LLM Speed Analysis: Is 28-35 tok/s Good for General Chatting? ## Understanding Token Generation Speed A generation speed of 28-35 tokens per second falls into the **good to very good** range for most general chatting and conversational AI applications. To put this in perspective, let me break down what this means in practical terms and how it compares to various benchmarks. ## Real-World Performance Context **Reading Speed Comparison:** - Average human reading speed: 200-300 words per minute (roughly 3-5 words per second) - Your LLM speed: 28-35 tokens per second - Since tokens ≈ 0.75 words on average, this translates to approximately 21-26 words per second This means your LLM is generating text about **5-8 times faster** than most humans can comfortably read, which creates an excellent user experience where responses appear to stream naturally without feeling sluggish. ## User Experience Benchmarks **Excellent (30+ tok/s):** Your upper range hits this threshold - Responses feel immediate and fluid - No noticeable waiting time for most queries - Suitable for real-time applications and interactive experiences **Good (20-30 tok/s):** Your lower range sits here - Still very responsive for general use - Slight streaming delay on longer responses - Perfectly adequate for most conversational AI applications **Acceptable (10-20 tok/s):** Below your range - Noticeable but not frustrating delays - Users might perceive slight sluggishness on complex queries ## Application-S…