Results from running Qwen3:27b q3 on Omen 16L / Intel 265f / 32GB / RTX 5060 Ti

303

As expected, it shows 20t/s exceeding 5050's 10t/s and handles larger models smoothly, outputting results

  • What I gave up on after running for an hour on M1 Max 32GB (7B, 14B ran but gave wrong answers, 27B was abandoned during execution)

  • On the 5050, the small model (14B Q2) took over 10 minutes

  • The 5060Ti 16GB gave an answer in just 2 minutes and handles subsequent questions lightly

At this level, it seems excellent for personal hobby development.

I've limited it to 150W for temperature management, and the graphics card can use up to 175W.

Indeed, I can't imagine local LLM without CUDA...

🪐 Planet Economics Engine Model Summary and Work Report


1️⃣ Model Overview: PlanetEconomicsEngine

PlanetEconomicsEngine is the core macroeconomic engine of the Flutter-based simulation game "OverStars". It calculates how planetary resources (population, gold, minerals, technology) change over time.

Key Features

  • Nonlinear Economics: Income tax/tariff friction modeling using the Laffer curve

  • Mineral Depletion Penalty: A 5x penalty is imposed on the treasury if minerals are scarce (triggers cascading bankruptcy)

  • Budget-Based Growth: Each ministry budget (social, environmental, education, military, economic) affects resource changes

3️⃣ Today's Work Summary

✅ Completed Tasks

  1. Model Analysis: Detailed analysis of the calculation logic in the PlanetEconomicsEngine.calculateNextState() function

  2. 10-Year Simulation: Resource change predictions up to 10 years with default initial values (population 1,000, gold 5,000, minerals 1,000)

  3. Calculation Documentation: Examples of formula application for each year and specific numerical values

  4. Difficulty-Specific Bankruptcy Scenario Design: Initial resource calculations for bankruptcy within 1/2/3 years based on difficulty levels (hard/normal/easy)

📊 Key Results

Item

Default Value (10-Year Simulation)

Population Growth Rate

~3.7% increase annually

Gold Decrease Rate

~876 annually (budget execution > tax revenue)

Mineral Decrease Rate

~150 annually (consumption > production)

Difficulty

Initial Gold

Initial Minerals

Expected Bankruptcy Time

Easy

2,500

400

~Year 3

Normal

1,500

250

~Year 2

Hard

800

120

~Year 1

4️⃣ Key Insights

  • Laffer Curve Effect: Raising tax rates indiscriminately actually decreases tax revenue (reduces work incentive)

  • Mineral Depletion Chain Reaction: When minerals are scarce, a 5x penalty is imposed on the treasury, accelerating bankruptcy

  • Importance of Balanced Budget: If military/administrative costs exceed tax revenue, continuous gold loss occurs

⚠️ Log Analysis Insights

Based on the logs and system specs (estimated HP Omen 16L, RTX 5060 Ti) you provided, I'll summarize the work results when running a large language model (Qwen 3 27B Q3, 32k Context) in that environment.

1. Core Performance Metrics (Speed & Throughput)

Despite running a heavy model of 27B (27 billion parameters) on hardware of this class (5060 Ti), the inference itself shows relatively stable speeds.

  • Text Generation Speed (Decoding): Approximately 21-23 tokens per second (tg = 21 ~ 23 t/s)

    • Faster than human reading speed, making it sufficient for real-time conversations or work.

  • Prompt Processing Speed (Pre-fill): Approximately 600-900 tokens per second

    • When the context is not very long (on cache hit), it reads the input conversation context very quickly.

2. Main Bottleneck: Cache Invalidation

This is the most prominent negative factor in the logs. When conversations get longer (task 147, task 5458, etc.), a phenomenon occurs where existing conversation memory (KV Cache) is completely discarded along with the following warning.

forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory...)

💡 Why Does It Happen?

Because of special memory techniques like Sliding Window Attention (SWA) used by the Qwen 3 model, when conversation content changes, the llama.cpp engine judges that existing context checkpoints are mathematically incompatible and destroys the cache.

⚠️ Resulting Damage (Lag Spike)

When the cache breaks, the prompt must be recalculated from the beginning when conversations get longer. Because of this, a severe delay (Lag) occurs where the system freezes for about 28-40 seconds before the first token appears before the response.

3. Memory and Resource Pressure

  • KV Cache Capacity: An additional ~4GB (3992.385 MiB) of VRAM/RAM is consumed just for remembering conversation content (prompts).

  • With the 32k context limit (n_ctx_slot = 32768) being fully utilized, 150MB checkpoints continue to be generated as conversations progress, putting memory pressure.

🛠️ Summary and Recommended Optimization Measures

Conclusion: "In the RTX 5060 Ti environment, the Qwen 3 27B model works excellently at about 22 tokens per second, but when conversations get longer, there are occasional outages of 30 seconds or more due to cache destruction."

To improve this, I recommend the following measures.

  1. Engine Update: The cache bug mentioned in the logs is a lightweight architecture issue being recently fixed by the llama.cpp open-source community. Update Ollama or your backend to the latest version.

  2. Context Size Reduction: If 32k context is not essential, reduce the setting to 16k (16384) or 8k (8192). The frequency of cache breaks and VRAM pressure will be dramatically reduced, making it much more comfortable.

로그인한 회원만 댓글 등록이 가능합니다.

개발한당

KR | ID | EN
  • IDR
  • KOR
7.52 ▼ -0.01

2026.10.09 KEB 하나은행 고시회차 3530회

다가오는 한인 행사일정

  • 등록 된 일정이 없어요!