Finding the Best CPU-Only LLM for Short Title Generation: Qwen 3.5 (0.8b 4b 9b) Benchmarks

avatar

Can you reliably run Small Language Models (SLMs) locally on a standard CPU without a dedicated graphics card? If you're looking for an efficient local model to handle specific tasks like post title generation, parameter size and inference speed are critical factors to balance.

In this experiment, I tested different model sizes from the Qwen family using Ollama directly in the terminal, completely relying on CPU processing (no dedicated GPU).

The goal: Generate a catchy post title under 40 characters based on a real fitness updates post sourced from Hive.


The Sample Input Text

Here is the original text passed to the models:

"Hello everyone, I hope everything is going well in your life and you are having a wonderful day. I woke up early this morning; the weather was clear and nice. Today I managed to exceed 7K. According to Google Fit, I walked 8400 steps, 6 km, and burned 2400 Cal. Take care of your health and enjoy your life in a healthy way.

Stay healthy and with love!
Thank you for reading"


Ollama Setup & Configuration

To keep tests fair and optimized for title generation without internal chain-of-thought overhead, I used the following Ollama parameters:

/set system Specialist in creating titles of up to 40 characters for texts.
/set nothink
/set parameter num_ctx 32000
/set nohistory
/set verbose

Test Results & Performance Metrics

1. Qwen (0.8B Parameters)

Generated Titles:

  1. Your morning walk was fantastic! The sunshine is perfect today! Keep shining bright on the trail ahead 🌤💪 #MorningGoals #HealthyWalk
  2. Today's Walk: From Clear Morning to Overcoming Obstacles 🌿💪✨ #HealthDay #7KGoal #WalkingTolerance #FitnessInspo #LifestyleGoals ❤😊🥗
  3. You're doing great – keep it up! 💪✨ #Fitness #HealthGoals #WalksStepUp💓🏃‍♀K

Performance Metrics:

  • Total Duration: 1.83s
  • Load Duration: 383.94ms
  • Prompt Eval Rate: 274.43 tokens/s
  • Eval Rate (Generation): 34.33 tokens/s

Analysis:
The 0.8B model is blazing fast on CPU (~34 tokens/s generation rate). However, it completely failed to adhere to the system prompt constraint of keeping titles under 40 characters. Instead, it produced long, social-media-style captions loaded with hashtags and emojis.


2. Qwen (4B Parameters)

Generated Titles:

  1. Exceeded 7K: Walked 8km & Burned 2400 Cals 🚶‍♂ #DailySteps
  2. Exceeded 7K: Walked 6km & Burned 2400 Calories 🚶‍♂ #FitnessJourney
  3. Exceeded 7K today: 6km walked & 8400 steps achieved 🚶‍♂e

Performance Metrics:

  • Total Duration: 13.94s
  • Load Duration: 8.63s (Model loading overhead)
  • Prompt Eval Rate: 57.91 tokens/s
  • Eval Rate (Generation): 11.30 tokens/s

Analysis:
The 4B model showed a significant improvement in capturing key data points from the text (e.g., "Exceeded 7K", "6km walked", "2400 Cals"). While closer to the goal, it still leaned heavily on hashtags and slight formatting artifacts at the end of output. Generation speed dropped to ~11.3 tokens/s on CPU.


3. Qwen (9B Parameters)

Generated Titles:

  1. Exceeded 7K: A Fresh Morning & Vitality
  2. Waking Early: Exceeded 7K Steps Today
  3. Exceeding 7K Steps: A Healthy Morning Routine

Performance Metrics:

  • Total Duration: 2.37s
  • Load Duration: 376.02ms (Model pre-loaded)
  • Prompt Eval Rate: 500.77 tokens/s
  • Eval Rate (Generation): 7.49 tokens/s

Analysis:
The 9B model delivered the best quality by far. It strictly respected the system prompt (< 40 characters), omitted unnecessary hashtags/emojis, and produced clean, professional blog titles. Although generation speed slowed down to ~7.5 tokens/s, it remains very usable for local workflow automation.


Comparison Summary

Metric / FeatureQwen 0.8BQwen 4BQwen 9B
Output QualityLow (Ignored length constraint)Medium (Good data extraction)High (Concise, accurate)
Instruction Following❌ Poor⚠️ Moderate✓ Excellent
Eval Speed (Tokens/s)34.3311.307.49
Best Use CaseUltra-fast draftingBalanced tasksFinal title generation

Conclusion

If your goal is high-quality, constrained text generation (like creating accurate titles under strict character limits), the Qwen 9B model is the clear winner. Despite running purely on CPU at ~7.5 tokens per second, the quality and compliance with system prompts far outweigh the faster generation speeds of smaller parameters.



0
0
0.000
0 comments