<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://thakicloud.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://thakicloud.github.io/" rel="alternate" type="text/html" /><updated>2026-07-20T08:47:24+09:00</updated><id>https://thakicloud.github.io/feed.xml</id><title type="html">Thaki Cloud Tech Blog | ThakiCloud | 다키클라우드 기술 블로그</title><subtitle>Thaki Cloud (ThakiCloud, 다키클라우드, thaki cloud, THAKI CLOUD, ثاكي كلاود)는 AI/ML Engineering, LLMOps, DevOps 분야의 최신 기술과 실무 경험을 공유하는 전문 기술 블로그입니다. 머신러닝 모델 운영, 쿠버네티스, 클라우드 인프라, AI 엔지니어링 커리어, 인공지능 기술 블로그, 다키클라우드 개발 팀의 깊이 있는 인사이트를 제공합니다. مدونة تقنية متخصصة في هندسة الذكاء الاصطناعي والحوسبة السحابية.</subtitle><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><entry xml:lang="ar"><title type="html">تشريح Kimi Code CLI: كيف يستحوذ وكيل الطرفية مفتوح المصدر على المحررات عبر ACP</title><link href="https://thakicloud.github.io/ar/agentops/kimi-code-cli-acp-open-source-agent/" rel="alternate" type="text/html" title="تشريح Kimi Code CLI: كيف يستحوذ وكيل الطرفية مفتوح المصدر على المحررات عبر ACP" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/ar/agentops/kimi-code-cli-acp-open-source-agent</id><content type="html" xml:base="https://thakicloud.github.io/ar/agentops/kimi-code-cli-acp-open-source-agent/"><![CDATA[<p>في الأسبوع الماضي، تصدرت Moonshot AI قوائم الترتيب في البرمجة بعد إطلاق نموذجها المفتوح الأوزان Kimi K3. لكن ما رافق هذا الإطلاق بهدوء كان أداة أقرب إلى سير عمل المطورين من النموذج نفسه، وهي Kimi Code CLI، وكيل برمجة طرفي مفتوح المصدر أطلقته Moonshot برخصة MIT. تداول مستخدمو LinkedIn تعريفاً يقول إن هذه الأداة تقدم ميزات غير موجودة في Claude Code. لم ننقل هذه العبارة كما هي، بل تحققنا مباشرة من المستودع الرسمي والوثائق. والخلاصة أن نصف هذا التعريف صحيح ونصفه الآخر مبالغ فيه. والنقطة الأكثر إثارة للاهتمام كانت في موضع لم تبرزه مواد الترويج.</p>

<p>يستعرض هذا المقال ماهية Kimi Code CLI، وما تقدمه فعلياً، ولماذا تستحق المتابعة من منظورنا كمشغّلين لمنصة ذكاء اصطناعي قائمة على K8s. وقد خصصنا مساحة واسعة لشرح لماذا يُعد معيار Agent Client Protocol المفتوح قطعة قد تغير قواعد اللعبة في منظومة الوكلاء.</p>

<h2 id="ما-هو-kimi-code-cli">ما هو Kimi Code CLI</h2>

<p>Kimi Code CLI أداة برمجة عاملة بأسلوب الوكيل تعمل من الطرفية، وتنتمي إلى نفس فئة Claude Code وGemini CLI وCodex CLI. المستودع الرسمي هو <a href="https://github.com/MoonshotAI/kimi-code">MoonshotAI/kimi-code</a>، وقد تطور من المشروع السابق <a href="https://github.com/MoonshotAI/kimi-cli">MoonshotAI/kimi-cli</a> مع الحفاظ على استمرارية الجلسات والإعدادات القديمة. كلا المستودعين رسميان من Moonshot، وينبغي الحذر من الخلط بينهما وبين مشاريع طرف ثالث تحمل أسماء مشابهة.</p>

<p>وهنا يجب توضيح تصحيح أول. الاسم الرسمي لهذه الأداة ليس “CLI مخصصة لـ Kimi K3” بل Kimi Code CLI. فهي أداة غير مقيدة بنموذج واحد، وتستخدم افتراضياً نموذج Moonshot المتخصص في البرمجة Kimi K2.7 Code، لكن يمكن عبر الإعدادات التحول إلى نماذج أخرى بما فيها K3. أي أن K3 هو أحد النماذج المتعددة التي يمكن ربطها بهذه الأداة، وليست الأداة مصممة حصراً له. أما K3 نفسه فهو نموذج MoE مفتوح بحجم 2.8 تريليون معلمة أطلقته Moonshot في 16 يوليو 2026، ويعتمد على Kimi Delta Attention مع نافذة سياق تصل إلى مليون رمز. وقد غطت هذا الإطلاق وسائل إعلام رئيسية مثل CNBC وBloomberg وForbes.</p>

<p>من المفيد رسم الصورة الكاملة أولاً لتترابط التفاصيل لاحقاً بسهولة أكبر. جوهر الأمر هو الدور المزدوج الذي تلعبه Kimi Code CLI: فمن جهة تتصل بالأدوات والبيانات كعميل MCP، ومن جهة أخرى تتصل بالمحررات كخادم ACP.</p>

<pre><code class="language-mermaid">flowchart TB
    subgraph EDITOR["محرر المطور (عميل ACP)"]
        ZED["Zed"]
        JB["عائلة JetBrains"]
        VSC["VS Code / Neovim"]
    end
    ACP["Agent Client Protocol&lt;br/&gt;JSON-RPC over stdio"]
    subgraph CLI["Kimi Code CLI (نواة الوكيل)"]
        MAIN["الوكيل الرئيسي&lt;br/&gt;يحافظ على سجل المحادثة"]
        SUB["الوكلاء الفرعيون&lt;br/&gt;coder · explore · plan&lt;br/&gt;لكل منهم سياق معزول"]
    end
    MODEL["طبقة النموذج&lt;br/&gt;Kimi K2.7 Code / K3&lt;br/&gt;أو نقطة نهاية متوافقة مع OpenAI"]
    subgraph MCP["خوادم MCP (الأدوات والبيانات)"]
        T1["Context7"]
        T2["Chrome DevTools"]
        T3["موصلات داخلية"]
    end

    EDITOR --&gt; ACP
    ACP --&gt;|kimi acp| MAIN
    MAIN --&gt; SUB
    MAIN --&gt;|طلب استدلال| MODEL
    SUB --&gt;|طلب استدلال| MODEL
    MAIN --&gt;|استدعاء أداة| MCP
</code></pre>

<h2 id="الوكلاء-الفرعيون-تقسيم-السياق-للحفاظ-على-نظافة-الوكيل-الرئيسي">الوكلاء الفرعيون: تقسيم السياق للحفاظ على نظافة الوكيل الرئيسي</h2>

<p>توفر Kimi Code CLI ثلاثة أنواع من الوكلاء الفرعيين المدمجين. الوكيل coder هو المسؤول الهندسي العام الذي يقرأ الملفات ويكتبها وينفذ الأوامر لتطبيق التغييرات الفعلية. الوكيل explore مخصص للاستكشاف، إذ يتصفح قاعدة الشيفرة للقراءة فقط. أما الوكيل plan فيقتصر عمله على تقديم خطط التنفيذ وتصاميم البنية دون تنفيذ أي أوامر في الصدفة. هذا التقسيم موثق في الصفحة الرسمية <a href="https://moonshotai.github.io/kimi-code/en/customization/agents.html">Agents and Sub-Agents</a>.</p>

<p>الجوهر هنا ليس الأسماء بل عزل السياق. يمتلك كل وكيل فرعي نافذة سياق مستقلة تماماً، ولا يرى سوى وصف المهمة الذي يمرره الوكيل الرئيسي صراحة. لا يُكشف سجل محادثة الوكيل الرئيسي للوكلاء الفرعيين، كما أن سجلات الاستدلال الوسيط واستدعاءات الأدوات التي ينفذها الوكيل الفرعي لا تختلط بسجل الوكيل الرئيسي، إذ يعيد الوكيل الفرعي النتيجة النهائية فقط. ولهذا السبب يظل السياق الرئيسي رفيعاً ولا يتضخم بالسجلات في الجلسات الطويلة. كما تدعم الأداة التنفيذ في الخلفية والتنفيذ المتوازي، بحيث يمكن تشغيل عدة مهام استكشاف في آن واحد وتعود النتائج تلقائياً عند الانتهاء.</p>

<p>هذا النمط ليس غريباً علينا. فحتى إطار التنسيق الداخلي الذي يشغّل هذه المدونة يفوّض مهام الاستكشاف إلى وكلاء فرعيين منخفضي التكلفة، ولا يستعيد سوى الملخصات لحماية السياق الرئيسي. مبدأ أن نظافة السياق تعني في الوقت نفسه التكلفة والجودة يبقى واحداً بغض النظر عن الأداة المستخدمة.</p>

<h2 id="mcp-تجربة-إعداد-دون-تعديل-json-يدوياً">MCP: تجربة إعداد دون تعديل JSON يدوياً</h2>

<p>تُدار عمليات ربط Model Context Protocol عبر مسارين. الأول هو الأوامر الفرعية لسطر الأوامر، حيث تُدار الخوادم عبر kimi mcp add وkimi mcp list وkimi mcp remove وkimi mcp authorize. على سبيل المثال يمكن ربط خادم بحث في الوثائق عبر نقل HTTP، أو ربط خادم أتمتة متصفح عبر نقل stdio.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># نقل HTTP (يدعم خيار OAuth)</span>
kimi mcp add <span class="nt">--transport</span> http context7 https://mcp.context7.com/mcp

<span class="c"># ربط عملية محلية عبر نقل stdio</span>
kimi mcp add <span class="nt">--transport</span> stdio chrome-devtools <span class="nt">--</span> npx chrome-devtools-mcp@latest
</code></pre></div></div>

<p>أما المسار الثاني فهو الأمر التفاعلي بشرطة مائلة /mcp-config الذي يُستخدم داخل واجهة TUI، ويتيح إضافة الخوادم وتعديلها والمصادقة عليها دون تحرير ملف إعدادات JSON مباشرة. ويعرض الأمر /mcp قائمة الخوادم المتصلة حالياً والأدوات المحمّلة. والجزء الذي أبرزه تعريف LinkedIn، وهو أنه لا حاجة لتعديل JSON مباشرة، صحيح فعلاً. لكن هذه الميزة بحد ذاتها ليست غائبة عن Claude Code، وسنعود إلى هذه النقطة لاحقاً. الوثائق ذات الصلة موجودة في <a href="https://moonshotai.github.io/kimi-cli/en/customization/mcp.html">إعداد MCP</a>.</p>

<h2 id="agent-client-protocol-أهم-قطعة-في-هذه-الأداة">Agent Client Protocol: أهم قطعة في هذه الأداة</h2>

<p>هذا هو الجزء الأكثر إثارة للاهتمام في هذا المقال. Agent Client Protocol، ويُختصر بـ ACP، هو معيار مفتوح صممه فريق محرر Zed. يعمل برخصة Apache، ويتبادل الرسائل عبر JSON-RPC 2.0 فوق stdio. تُشغّل المحررات الوكيل كعملية فرعية وتتواصل معه عبر المدخلات والمخرجات القياسية، وآلية النقل نفسها مطابقة لبروتوكول خادم اللغة.</p>

<p>التشبيه هنا يساعد كثيراً على الفهم. قبل ظهور LSP، كان على كل محرر أن يبني تكاملاً منفصلاً لكل لغة برمجة. حوّل LSP هذه المسألة من مشكلة M ضرب N إلى مشكلة M زائد N، فما إن ينفذ محرر واحد المعيار حتى يستفيد من أي خادم لغة بغض النظر عمن صنعه. ويفعل ACP الشيء ذاته تماماً مع الوكلاء، فما إن ينفذ محرر واحد ACP حتى يتصل به أي وكيل بطريقة موحدة بغض النظر عن صانعه. يمكن الاطلاع على هذا المفهوم في <a href="https://zed.dev/acp">تعريف Zed بـ ACP</a> وفي <a href="https://blog.marcnuri.com/agent-client-protocol-acp-introduction">مقال مارك نوري التوضيحي</a>.</p>

<p>من السهل الخلط بينه وبين MCP، لكن الاتجاه معاكس تماماً. يتجه MCP من الوكيل نحو الأدوات والبيانات، وفي هذه الحالة يكون الوكيل عميل MCP. أما ACP فيتجه من المحرر نحو الوكيل، وهنا يكون الوكيل خادم ACP والمحرر عميل ACP. أي أن الوكيل نفسه يلعب دور عميل MCP من جهة، ودور خادم ACP من جهة أخرى في آن واحد. وهذا هو السبب الذي جعل الرسم البياني السابق يوضح هذا الدور المزدوج.</p>

<p>تدعم Kimi Code CLI هذا البروتوكول بشكل أصلي عبر الأمر الفرعي kimi acp دون الحاجة إلى أي تثبيت إضافي. يتصل بها Zed بشكل أصلي، بينما تتصل به JetBrains عبر إضافة، وقد ظهرت بالفعل عدة تكاملات مع محررات أخرى وفق سجل ACP الخاص بـ Zed. بذلك يستطيع المطور تشغيل جلسة Kimi دون مغادرة المحرر الذي اعتاد عليه.</p>

<h2 id="إدخال-الصور-والفيديو-إلى-أي-حد-فعلياً">إدخال الصور والفيديو، إلى أي حد فعلياً</h2>

<p>ذكر تعريف LinkedIn أنه يمكن تمرير لقطة الشاشة كما هي كمدخل. وهذا يحتاج إلى تصحيح. فالميزة التي تبرزها Moonshot فعلياً ليست لقطة شاشة ثابتة بل إدخال مقطع فيديو مسجّل للشاشة. يذكر وصف المستودع أنه عند إسقاط تسجيل شاشة أو مقطع عرض توضيحي في المحادثة، يستطيع الوكيل مشاهدة وفهم السلوك الذي يصعب شرحه بالكلام مباشرة. وبالطبع تدعم نافذة الإدخال في سطر الأوامر لصق الصور أيضاً، إذ إن النموذج الافتراضي Kimi K2.7 Code هو نموذج متعدد الوسائط أصلي مزوّد بمُرمّز رؤية MoonViT بحجم 400 مليون معلمة، يستقبل النص والصور والفيديو معاً. غير أنه عند ربط نموذج مخصص، يجب تحديد دعم الصور صراحة ضمن modalities الخاصة بذلك النموذج ليعمل بشكل صحيح. وخلاصة القول إن إدخال الصور متاح فعلاً، لكن ما يُروَّج له كفارق حقيقي هو إدخال الفيديو، وتعبير لقطة شاشة غير دقيق تماماً.</p>

<h2 id="التثبيت-فعلياً-في-ثلاث-خطوات">التثبيت فعلياً في ثلاث خطوات</h2>

<p>تدفق التثبيت بسيط فعلاً كما ورد في التعريف. الأوامر أدناه مستندة إلى <a href="https://moonshotai.github.io/kimi-cli/en/guides/getting-started.html">دليل البدء الرسمي</a>، ولم نترك سجل تنفيذ مباشر لأن صندوق الاختبار الداخلي لدينا لا يملك صلاحية الوصول إلى نطاق التوزيع المعني. لذلك لم ننتج أي أرقام قياس أداء، واكتفينا بنقل الأوامر الموثقة فقط.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># 1) تشغيل سكربت التثبيت (يثبّت uv في الوقت نفسه)</span>
curl <span class="nt">-LsSf</span> https://code.kimi.com/install.sh | bash

<span class="c"># 2) التشغيل داخل دليل المشروع</span>
kimi

<span class="c"># 3) إعداد المصادقة</span>
/login
</code></pre></div></div>

<p>بالنسبة لنظام macOS تتوفر brew install kimi-code، وبالنسبة لويندوز يتوفر أيضاً سكربت PowerShell. للتطوير من الشيفرة المصدرية يلزم Node بإصدار 24.15 أو أحدث مع pnpm. ولأن الرخصة MIT، فإن القيود قليلة على قراءة الشيفرة وعمل fork لها وتوزيعها داخل المؤسسة.</p>

<h2 id="انفتاح-النماذج-ومزودي-الخدمة">انفتاح النماذج ومزودي الخدمة</h2>

<p>يبلغ أقصى طول للسياق في عائلة K2.6 نحو 256 ألف رمز، بينما يصل في K3 وفق مواد تسويق Moonshot إلى مليون رمز. لكن الأهم من ذلك هو انفتاح مزودي الخدمة. ففي ملف ~/.kimi-code/config.toml يمكن تسجيل عدة مزودين في آن واحد، من نقاط نهاية متوافقة مع OpenAI، إلى مفاتيح Anthropic API، وصولاً إلى Google GenAI أو Vertex AI. وهذا يعني أن الأداة غير مقيدة بنموذج واحد بعينه. كما تعالج تلقائياً حقل reasoning_content الخاص بنماذج الاستدلال من أطراف ثالثة. الوثائق ذات الصلة في <a href="https://moonshotai.github.io/kimi-cli/en/configuration/providers.html">Providers and models</a>.</p>

<h2 id="هل-هي-ميزات-غير-موجودة-في-claude-code-مقارنة-صريحة">هل هي ميزات غير موجودة في Claude Code: مقارنة صريحة</h2>

<p>أكثر العبارات انتشاراً في التعريف كانت أنها تقدم ميزات غير موجودة في Claude Code. وبعد التحقق تبين أن هذا الإطار في معظمه مبالغ فيه.</p>

<p>فالوكلاء الفرعيون وعزل السياق يقدمهما Claude Code أيضاً بالطريقة نفسها عبر ميزة الوكلاء الفرعيين. وMCP مدعوم أصلاً بنضج في Claude Code عبر نقل stdio وSSE وHTTP. كما أن لصق الصور موجود فيه أيضاً. هذه العناصر الثلاثة إذن ليست فوارق حقيقية.</p>

<p>الفارق الحقيقي يكمن في نقطتين. الأولى هي طريقة دعم ACP. فـ Kimi Code CLI تدمج ACP كميزة أساسية من الدرجة الأولى داخل الأداة نفسها عبر الأمر الفرعي kimi acp. أما Claude Code فيتصل بها عبر حزمة محول منفصلة صنعها Zed، وما تزال في مرحلة تجريبية. من منظور المستخدم، الأولى تعمل فور تفعيل الأداة، بينما الثانية تتطلب إضافة جسر إضافي. النقطة الثانية هي انفتاح النماذج. فـ Kimi مفتوح على سلسلة K مفتوحة الأوزان مع إمكانية التحول بين عدة مزودين، بينما يقتصر Claude Code على نماذج Anthropic حصرياً. ومن هذه النقطة يتفرع فارق ثالث يتعلق بإمكانية الاستضافة الذاتية. فبما أن Kimi أداة مفتوحة المصدر مع نموذج مفتوح الأوزان، يمكن تشغيله داخل المؤسسة، بينما تظل Claude Code أداة مفتوحة لكن نموذجها متاح عبر واجهة API فقط. يمكن الاطلاع على الدليل ذي الصلة في <a href="https://zed.dev/blog/claude-code-via-acp">مقال Zed حول Claude Code عبر ACP التجريبي</a>.</p>

<h2 id="دلالات-على-منتجات-thakicloud">دلالات على منتجات ThakiCloud</h2>

<p>يمس هذا الموضوع أداة وكيل من جهة، ومحور بنية تحتية يتعلق بالنماذج المفتوحة والاستضافة داخل المؤسسة من جهة أخرى. لذلك نستخدم العدستين معاً.</p>

<p>من عدسة Paxis، تتداخل بنية Kimi Code CLI إلى حد كبير مع اتجاه تصميم منتجنا. فPaxis هو مستوى التحكم الخاص بـ ThakiCloud لسحابة أصيلة الوكلاء (Agent-Native Cloud)، ويتعامل مع المهارات والأدوات والسياسات وسجلات التدقيق كموارد من الدرجة الأولى. والطريقة التي يعمل بها الوكلاء الفرعيون coder وexplore وplan لدى Kimi بالتوازي وفي سياقات معزولة، تشترك في الفلسفة نفسها مع طريقة عمل حاضنة المهارات في Paxis، التي تختار من بين أكثر من 960 مهارة باستخدام BM25 وتنفذها في صناديق اختبار معزولة. وACP بشكل خاص، بوصفه معياراً محايداً تجاه المزودين، يمثل فرصة مباشرة لـ Paxis. فأي وكيل ننشره، بما في ذلك الوكلاء المزودة بنماذج خضعت لضبط دقيق خاص بنا، يمكنه إذا نفذ ACP أن يتصل بمحررات المطورين لدى العملاء مثل Zed أو JetBrains بطريقة موحدة. وهذا المزيج من معيارين، MCP للاتصال بالبيانات وACP للاتصال بالمحررات، يمثل بالضبط الصورة التكاملية التي نتجه إليها.</p>

<p>ومن عدسة ai-platform، الانفتاح يعني حرية النشر مباشرة. فبوضع سلسلة K مفتوحة الأوزان فوق جدولة GPU عبر Kueue وخدمة vLLM في عنقودنا، وتوجيه الأداة نحو نقطة نهاية داخلية، يمكن بناء وكيل برمجة داخلي دون الاعتماد على API خارجي أو إخراج البيانات إلى الخارج. وهذا ينسجم مع متطلبات الأمن الخاصة بالاستضافة داخل المؤسسة في قطاعي المال والقطاع العام حيث لا يجوز خروج الشيفرة إلى الخارج، وكذلك مع متطلبات جهات مثل NIS. وقد سبق أن تناولنا في مقالات سابقة فكرة أنه كلما أصبحت القدرات شائعة ورخيصة، فإن ما تدفعه الشركات فعلياً هو بيئة تنفيذ محكومة. وأهمية Kimi Code CLI تكمن في أنها فتحت طبقة التنفيذ هذه كمصدر مفتوح.</p>

<h2 id="القيود-والاعتراضات">القيود والاعتراضات</h2>

<p>هناك عدة نقاط ينبغي النظر إليها بموضوعية. أولاً، أسماء المحركات الداخلية أو هياكل الطبقات التي تُذكر في تحليلات طرف ثالث معمقة لا ترد في الوثائق الرسمية، وقد تكون نتيجة هندسة عكسية، لذا من الأسلم الاستناد إلى الوثائق الرسمية قبل اعتبارها حقائق. ثانياً، توجد تقارير من المجتمع تفيد بأن مسار ACP يقدم جودة استجابة أفضل من طرق الاتصال الأخرى، لكنها انطباعات وليست قياسات أداء موثقة، أي أنها ليست أرقاماً محققة. ثالثاً، حتى مع كون النموذج مفتوح الأوزان، فإن خدمة نموذج بحجم 2.8 تريليون معلمة فعلياً داخل المؤسسة تتطلب موارد GPU كبيرة، والانفتاح لا يعني بالضرورة سهولة الاستضافة الذاتية، إذ يظل مسار API خياراً واقعياً للفرق الصغيرة. رابعاً، قد يتفوق نضج واستقرار منظومة الأدوات لدى Claude Code أو Codex CLI. وكون الأداة مفتوحة المصدر لا يعني بالضرورة أنها جاهزة للإنتاج.</p>

<p>ومع ذلك، فإن الاتجاه نحو ارتباط رخو بين الوكلاء والمحررات فوق معايير مفتوحة هو تيار واضح. فعالم يستطيع فيه المطور تبديل النموذج والمحرر كل على حدة دون التقيد بأداة سطر أوامر من مزود معين، هو عالم أكثر فائدة للمطورين. وتُعد Kimi Code CLI إحدى القطع التي تُقرّب هذا العالم.</p>

<h2 id="المصادر">المصادر</h2>

<ul>
  <li><a href="https://github.com/MoonshotAI/kimi-code">MoonshotAI/kimi-code (المستودع الرسمي)</a></li>
  <li><a href="https://github.com/MoonshotAI/kimi-cli">MoonshotAI/kimi-cli (المستودع السابق)</a></li>
  <li><a href="https://moonshotai.github.io/kimi-cli/en/guides/getting-started.html">دليل بدء Kimi Code CLI</a></li>
  <li><a href="https://moonshotai.github.io/kimi-code/en/customization/agents.html">وثائق Agents and Sub-Agents</a></li>
  <li><a href="https://moonshotai.github.io/kimi-cli/en/customization/mcp.html">وثائق إعداد MCP</a></li>
  <li><a href="https://zed.dev/acp">Zed - Agent Client Protocol</a></li>
  <li><a href="https://blog.marcnuri.com/agent-client-protocol-acp-introduction">ACP: The LSP for AI Coding Agents</a></li>
  <li><a href="https://zed.dev/blog/claude-code-via-acp">Zed - Claude Code via ACP (تجريبي)</a></li>
  <li><a href="https://www.marktechpost.com/2026/07/16/moonshot-ai-releases-kimi-k3-a-2-8-trillion-parameter-open-moe-model-with-kimi-delta-attention-and-1m-context/">MarkTechPost - إطلاق Kimi K3</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="agentops" /><category term="agentops" /><category term="kimi" /><category term="moonshot" /><category term="coding-agent" /><category term="mcp" /><category term="agent-client-protocol" /><category term="paxis" /><category term="thakicloud" /><summary type="html"><![CDATA[نحلل أداة سطر الأوامر المفتوحة المصدر للبرمجة التي أطلقتها Moonshot AI إلى جانب Kimi K3، استناداً إلى الوثائق الرسمية والمستودع الفعلي. من الوكلاء الفرعيين coder وexplore وplan، مروراً بإعداد MCP التفاعلي، وصولاً إلى الدعم الأصلي لـ Agent Client Protocol الذي يمثل الفارق الحقيقي، نتحقق إلى أي مدى تصح عبارة الترويج القائلة إن هذه ميزات غير موجودة في Claude Code.]]></summary></entry><entry xml:lang="ar"><title type="html">فاتورة بقيمة 25 مليار وون تطرح سؤالا: لماذا تكلفة وكلاء الذكاء الاصطناعي غير مرئية؟</title><link href="https://thakicloud.github.io/ar/llmops/agent-cost-observability-billing-crisis/" rel="alternate" type="text/html" title="فاتورة بقيمة 25 مليار وون تطرح سؤالا: لماذا تكلفة وكلاء الذكاء الاصطناعي غير مرئية؟" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/ar/llmops/agent-cost-observability-billing-crisis</id><content type="html" xml:base="https://thakicloud.github.io/ar/llmops/agent-cost-observability-billing-crisis/"><![CDATA[<p>كُتب هذا المقال لمهندسي المنصات والبنية التحتية الذين يخططون لإدخال Claude Code أو وكلاء الذكاء الاصطناعي إلى مؤسساتهم، وللمسؤولين الماليين ومسؤولي المشتريات الذين سيضطرون لتفسير فاتورة الذكاء الاصطناعي في الشهر القادم. ولنبدأ بالخلاصة: أخبار تكلفة الذكاء الاصطناعي التي توالت خلال الشهر الأخير لا تدور حول أن “الذكاء الاصطناعي مكلف”. المشكلة الحقيقية هي أن <strong>الفاتورة لا تفسر ما الذي تمثله</strong>. ففي بنية الوكلاء، حيث يتحول طلب مستخدم واحد إلى عشرات أو مئات من استدعاءات النموذج وتنفيذ الأدوات، ثم إلى إعادة محاولة تلقائية عند الفشل، لا يمكن للمبلغ النهائي وحده أن يكشف أين تسربت الأموال داخل أي حلقة تنفيذ. نرى أن هذه الفجوة في المراقبة هي بالضبط جوهر الألم الذي يعيشه السوق الآن.</p>

<h2 id="نظرة-عامة">نظرة عامة</h2>

<p>يمكن تلخيص أخبار الفترة من أواخر يونيو حتى يوليو 2026 في سطر واحد: الاستخدام العشوائي للنماذج المتطورة (frontier models) في كل مهمة يجعل التكلفة يصعب تحملها، بل ويصعب حتى تتبع مصدرها. وقد ظهرت الحادثة بشكل درامي. فقد جرت محاولة تحصيل مبلغ يقارب 2.5 مليار وون من مستخدم محلي واحد، ثم محاولة تحصيل نحو 25 مليار وون لاحقا. ولحسن الحظ لم يُسحب أي مبلغ فعليا بسبب تجاوز حد البطاقة، غير أن تكرار وصول مبلغ غير طبيعي إلى مرحلة طلب موافقة البطاقة، وليس مجرد خطأ عرض، هو ما يمنح الحادثة وزنها المختلف.</p>

<p>وفي الفترة نفسها توالت أخبار من مستويات أخرى. فقد راجعت شركة تدقيق تكاليف الذكاء الاصطناعي فواتير 60 شركة وادّعت وجود فوترة زائدة كبيرة، وبدأت عدة شركات كبرى بتوزيع استخدام النماذج المتطورة على نماذج أرخص بحسب طبيعة المهمة، كما ورد أن شركات أمريكية وأوروبية انتقلت إلى نماذج صينية مفتوحة الأوزان بدافع خفض التكلفة. والمثير أن مزودي النماذج أنفسهم بدؤوا بالرد بما يفيد أن “تشغيل أفضل النماذج أداءً لفترات طويلة في كل مهمة أمر غير مستدام”. وكانت أخبار متعددة الاتجاهات تشير إلى النقطة نفسها.</p>

<h2 id="ماذا-حدث-في-الشهر-الأخير">ماذا حدث في الشهر الأخير</h2>

<p>أول ما لفت الانتباه كان مشكلة موثوقية نظام الفوترة. وبحسب تقرير ZDNet Korea، بدأ المبلغ المطلوب من طالب جامعي محلي بنحو 1.66 مليون دولار، ثم تضخم إلى نحو 16.62 مليون دولار أي عشرة أضعاف. وأوضحت أنثروبيك لاحقا أن الخطأ كان في إعداد مبلغ الشحن التلقائي الذي ضُبط بشكل غير طبيعي المرتفع، إلا أن المستخدم أكد أنه لم يُفعّل خاصية الشحن التلقائي إطلاقا، ما يترك سبب نشوء هذا الإعداد غامضا حتى الآن. وتُظهر واقعة إرسال المستخدم أكثر من خمس عشرة رسالة إلى عدة أقسام دون تلقي رد آلي إلا بعد مرور أربعة أيام، فجوةً في منظومة الاستجابة أوضح من كونها مجرد خطأ تقني.</p>

<p>المشكلة الثانية كانت في قابلية مراقبة تكلفة الوكلاء. وبحسب تقارير AI Times و The Information، أعلنت شركة ناشئة متخصصة في تدقيق التكاليف تُدعى Vaudit أنها راجعت فواتير 60 شركة بقيمة إجمالية نحو 34 مليون دولار، وخلصت إلى أن نحو 1.7 مليون دولار منها كانت فوترة زائدة. وشكّل استخدام Claude Code جزءا كبيرا من نطاق المراجعة، وذُكرت شركات مثل باناسونيك وHP وهوندا ضمن العملاء. وحسب ادعاء الشركة، فإن الأنماط شملت تسجيل استخدام نموذج رخيص بأسعار نموذج أغلى، وفرض رسوم على مهام لم تُنجز، وتكرار إعادة المحاولة التلقائية بعد الخطأ فيما يُعرف بـ<strong>عاصفة إعادة المحاولة (retry storm)</strong>. وهنا ينبغي توضيح نقطتين. أولا، ردت أنثروبيك بأنها لا تفرض رسوما على الطلبات غير المكتملة أو الاستجابات الخاطئة، وأنه لا يوجد دليل واسع على فوترة زائدة. ثانيا، تحصل Vaudit على نسبة من مبالغ الاسترداد الناجحة، أي أنها شركة تدقيق تجارية، لذا يجب قراءة هذه الأرقام كنتيجة تحقيق من طرف واحد لا كتدقيق حسابي مستقل. وبعبارة أخرى، نحن الآن في مرحلة تصادم بين ادعاءات شركة التدقيق ونفي المزود.</p>

<p>المشكلة الثالثة كانت في استجابة السوق. وذكرت The Information أن الشركات بدأت بفصل المهام: النماذج الرخيصة للتصنيف والتلخيص والتحويل البسيط، والنماذج المتطورة للبرمجة المعقدة ومهام الوكلاء، والنماذج مفتوحة الأوزان أو المستضافة ذاتيا للمهام المتكررة عالية الحجم. ونقلت الفايننشال تايمز أن شركات مثل DoorDash وSiemens وAirbnb اعتمدت نماذج DeepSeek أو من عائلة Moonshot لخفض التكلفة. وفي تقرير لـ Business Insider، اعترف حتى مسؤولو المنصة في أنثروبيك بأن ما يُعرف بـ<strong>التقنية المعلوماتية الظل (shadow IT)</strong>، أي اعتماد كل قسم لأدواته الخاصة بشكل منفصل، أدى إلى تضخم تكاليف الذكاء الاصطناعي في بعض الشركات، لكنهم شددوا على أن الحل ليس إيقاف الاستخدام أو فرض سقف ميزانية موحد، بل اختيار النموذج بحسب المهمة وإدارة مركزية للتكلفة على مستوى المؤسسة. كما تغيرت سياسات التسعير نفسها مرارا: تعدّلت مرات عدة شمولية النماذج عالية الأداء الأحدث ضمن الاشتراك وتوقيت التحول إلى الدفع بحسب الاستخدام، وتكرر تمديد مواعيد انتهاء العروض الترويجية. بل ورد أن إتاحة Claude Fable 5 مجانا امتدت حتى 19 يوليو. وكانت صعوبة توقع تكلفة الشهر القادم أكبر إزعاج لمسؤولي المشتريات، أكثر حتى من مسألة الأداء.</p>

<h2 id="لماذا-لا-يمكن-مراقبة-تكلفة-الوكلاء">لماذا لا يمكن مراقبة تكلفة الوكلاء</h2>

<p>السبب المشترك الذي يجمع بين هذه المسارات الثلاثة من الأخبار هو في النهاية واحد. ففي أحمال عمل الوكلاء، اتسعت المسافة كثيرا بين ما يراه المستخدم وما تسجله الفاتورة. كانت استدعاءات API التقليدية طلبا واحدا مقابل استجابة واحدة وسطر تكلفة واحد. أما وكلاء البرمجة أو حزم تطوير الوكلاء (agent SDK)، فإن أمرا واحدا فيها يتوسع إلى وضع خطة وتنفيذ واستدعاء أدوات وتحرير ملفات وتحقق، ثم إعادة محاولة عند الفشل. يحدث هذا التوسع في مكان لا يراه المستخدم، بينما تسجل الفاتورة مجموعه الكلي في سطر واحد فقط.</p>

<pre><code class="language-mermaid">flowchart TB
    U["طلب مستخدم واحد"] --&gt; P["وضع خطة الوكيل"]
    P --&gt; L["حلقة التنفيذ"]
    L --&gt; T["استدعاء أدوات · استدعاء نموذج&lt;br/&gt;عشرات إلى مئات المرات"]
    T --&gt; R{"هل نجحت؟"}
    R --&gt;|"فشل"| RS["إعادة محاولة تلقائية&lt;br/&gt;(retry storm)"]
    RS --&gt; T
    R --&gt;|"نجاح"| ACC["توكن · كاش · tool call&lt;br/&gt;تجميع تراكمي"]
    ACC --&gt; INV["الفاتورة: مبلغ نهائي في سطر واحد"]
    INV -.فجوة المراقبة.-&gt; U
</code></pre>

<p>في هذه البنية، تقع معظم نقاط تسرب التكلفة خارج مجال رؤية المستخدم. فحلقة إعادة المحاولة تعمل بصمت وتضخم عدد الاستدعاءات، ويحدث تباين بين الاستخدام الفعلي للنموذج والتفاصيل النهائية للفاتورة عند المرور عبر مزود سحابي وسيط، وإذا اختل إعداد واحد مثل الشحن التلقائي فقد يتدفق مبلغ غير طبيعي حتى مرحلة طلب موافقة البطاقة. تبدو الأخبار الثلاثة وكأنها حوادث منفصلة، لكنها في الواقع أوجه مختلفة لفجوة المراقبة نفسها. لذلك لا يكفي مجرد وضع حد شهري لكل مستخدم. المطلوب هو طبقة قياس مركزية تلتقط تكلفة كل نموذج، وتوكنات كل جلسة، وتوكنات الكاش، وعدد استدعاءات الأدوات، وتكلفة الفشل وإعادة المحاولة، ومعدل الزيادة اليومي غير الطبيعي <strong>في اللحظة التي يحدث فيها الاستدعاء نفسه</strong>. فبلا مراقبة لا توجد سيطرة، وبلا سيطرة تبقى الفاتورة دائما وثيقة مفاجأة بعد وقوع الحدث.</p>

<h2 id="دلالات-التطبيق-على-منتجات-thakicloud">دلالات التطبيق على منتجات ThakiCloud</h2>

<p>هذه المشكلة هي النقطة التي يستهدفها منتجا ThakiCloud من زاويتين مختلفتين. ولأن منظور البنية التحتية ومنظور الوكلاء يكملان بعضهما، نستخدم في هذا الموضوع العدستين معا.</p>

<p><strong>عدسة ai-platform: الملكية هي الحل لأحمال العمل المتكررة.</strong> الاستنتاج الذي وصل إليه السوق واضح. فمعالجة حتى المهام السهلة بنماذج متطورة يجعل التكلفة غير محتملة، بينما تصبح استضافة نموذج مفتوح الأوزان ذاتيا خيارا اقتصاديا للمهام المتكررة عالية الحجم. ومنصة ai-platform من ThakiCloud هي بالضبط بنية تحتية للذكاء الاصطناعي وتعلم الآلة قائمة على K8s مصممة لهذه النقطة. فهي تستخدم Kueue لجدولة وحدات GPU في طابور وزيادة معدل استخدامها، وتستخدم vLLM لخدمة النماذج مفتوحة الأوزان، وتعزل الاستخدام بحسب كل قسم عبر عزل متعدد المستأجرين مع فوترة منفصلة. فإذا كانت واجهات برمجة التطبيقات القائمة على الدفع بحسب الاستخدام تنتج فواتير يصعب توقعها، فإن الاستضافة الذاتية تبني بنية تكلفة GPU ثابتة لا يتقلب سعرها حتى مع تزايد الاستخدام. وعلى عكس الأنظمة الخارجية القائمة على الدفع بحسب الاستخدام التي تتغير سياساتها باستمرار، يحوّل النشر الداخلي أو السيادي إمكانية التنبؤ بالتكلفة نفسها إلى أصل قائم بذاته. كما أن عدم خروج البيانات إلى الخارج يمثل قيمة إضافية للمؤسسات ذات متطلبات التنظيم والأمن المحلية العالية.</p>

<p><strong>عدسة Paxis: جعل كل تصرف للوكيل قابلا للتدقيق.</strong> كان جوهر فجوة المراقبة هو حلقة الوكيل، وهذا هو بالضبط المجال الذي تعالجه Paxis. فـPaxis هي مستوى التحكم في السحابة الأصيلة للوكلاء (Agent-Native Cloud) من ThakiCloud، وتعمل فوق ai-platform، وتتعامل مع المهارات (Skills) والأدوات (Tools) والسياسات (Policies) وسجلات التدقيق (Audit Logs) كموارد من الدرجة الأولى. فكل ما يخص أي مهارة استدعاها الوكيل وبأي أداة وكم مرة، وفي أي بيئة معزولة تم التنفيذ، يُسجَّل بالكامل في سجل التدقيق. وفي هذه البنية، بدلا من أن تُضخّم عاصفة إعادة المحاولة الفاتورة بصمت، تظهر حلقة إعادة المحاولة بوضوح في سجل التدقيق، وتوقف بوابات السياسة الاستدعاءات التي تتجاوز العتبة المحددة. فتصميم يختار من بين أكثر من 960 مهارة عبر خوارزمية BM25 وينفذها في بيئة معزولة، ويُمرّر كل تصرف عبر السياسة والتدقيق، هو إجابة بنيوية بالضبط على مشكلة صعوبة معرفة أين نشأت التكلفة من مجرد النظر إلى الفاتورة. فالخدمة منخفضة التكلفة (ai-platform) تجعل الوكلاء اقتصاديين، بينما تجعل المراقبة على مستوى كل تصرف (Paxis) هذه الجدوى الاقتصادية قابلة للتنبؤ. وهكذا تتكامل العدستان.</p>

<h2 id="الحدود-والرأي-المعاكس">الحدود والرأي المعاكس</h2>

<p>من أجل التوازن، نوضح الجانب المعاكس بجلاء. أولا، ليست النماذج المتطورة تبذيرا بالضرورة. فبحسب تقرير وول ستريت جورنال، ترى شركات مثل Shopify أن النماذج المتطورة، في مهام البرمجة المعقدة والوكلاء متعددة الخطوات، توفر وقت المهندسين بما يبرر سعرها المرتفع. في المقابل، تتوخى شركات مثل Spotify وTwilio الحذر في تقييم ما إذا كان التحسن الطفيف في الأداء يبرر التكلفة الإضافية. أي أن الجواب ليس “تخلَّ عن النماذج المتطورة”، بل “وزّع بحسب صعوبة المهمة”. كما أن الاستضافة الذاتية ليست حلا شاملا لكل شيء؛ فتخفيض المهام التي تتطلب أعلى مستوى من الاستدلال إلى نماذج مفتوحة الأوزان يؤدي إلى تراجع الجودة، ويخلق عبئا تشغيليا جديدا يشمل تشغيل وحدات GPU وتحديث النماذج وتصحيحات الأمان.</p>

<p>ثانيا، الأرقام المتعلقة بالفوترة الزائدة المذكورة في هذا المقال ليست حقائق مؤكدة. فادعاء Vaudit هو إعلان من شركة تدقيق تجارية، وقد نفته أنثروبيك، لذا فإن الأدق حاليا هو قراءة الموقف على أنه تصادم بين طرفين. وحادثة الفوترة بمبلغ 25 مليار وون كذلك لم يُسحب فيها أي مبلغ فعليا، ولم يُكشف بعد عن تفسير تقني لسبب نشوء إعداد الشحن التلقائي. والاستنتاج الذي نخلص إليه من هذه الأخبار لا يستهدف مزودا بعينه، بل هو مبدأ مفاده أنه في عصر الوكلاء، وأيا كان المزود المستخدَم، يجب على الجهة المستخدِمة نفسها أن تؤمّن مراقبة التكلفة وحوكمتها. فمسألة اختيار نموذج جيد ومسألة ضبط ذلك النموذج مسألتان منفصلتان، وما كشفته أخبار الشهر الأخير هو أن المسألة الثانية كانت فارغة طوال الوقت.</p>

<h2 id="المصادر">المصادر</h2>

<ul>
  <li><a href="https://zdnet.co.kr/view/?no=20260709165452">ZDNet Korea، “مستخدم محلي تلقى طلب دفع بقيمة 25 مليار وون… جدل حول خطأ فوترة أنثروبيك” (2026-07-09)</a></li>
  <li><a href="https://zdnet.co.kr/view/?no=20260716093004">ZDNet Korea، “أنثروبيك التي طالبت بـ25 مليار وون: تبين أنه خطأ في إعداد الشحن التلقائي” (2026-07-16)</a></li>
  <li><a href="https://www.aitimes.com/news/articleView.html?idxno=212155">AI Times، “أنثروبيك في جدل ‘الفوترة الزائدة للذكاء الاصطناعي’: تحصيل رسوم حتى على المهام الفاشلة” (2026-06)</a></li>
  <li><a href="https://www.theinformation.com/titv/fedld">The Information، تقرير عن سيطرة الشركات على تكلفة الذكاء الاصطناعي وتوزيع النماذج (2026-06-23)</a></li>
  <li><a href="https://www.ft.com/content/9c8ff45b-7c20-4c2e-93c9-c52339ffdcee">Financial Times، “Companies turn to Chinese AI models to cut costs” (2026-07)</a></li>
  <li><a href="https://www.businessinsider.com/anthropic-ai-costs-responses-routers-2026-7">Business Insider، “Anthropic Official Warns Against ‘Wrong’ AI Cost Response” (2026-07-15)</a></li>
  <li><a href="https://www.wsj.com/cio-journal/meet-the-companies-shelling-out-for-top-ai-models-e1fe3375">The Wall Street Journal، “Meet the Companies Shelling Out for Top AI Models” (2026-07)</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="llmops" /><category term="LLMOps" /><category term="FinOps" /><category term="تكلفة الوكلاء" /><category term="مراقبة التكلفة" /><category term="توجيه النماذج" /><category term="self-hosting" /><category term="Paxis" /><category term="بنية الذكاء الاصطناعي التحتية" /><summary type="html"><![CDATA[من حادثة فوترة بقيمة 25 مليار وون طالت مستخدما محليا إلى شبهات فوترة زائدة شملت 60 شركة، تشير أخبار تكلفة الذكاء الاصطناعي في الشهر الأخير إلى فجوة واحدة. في عصر الوكلاء حيث يتحول الطلب الواحد إلى مئات من استدعاءات النموذج، لم تعد الفاتورة تفسر ما الذي دُفع المال مقابله. هذا المقال يلخص كيفية سد فجوة المراقبة هذه.]]></summary></entry><entry xml:lang="ar"><title type="html">نموذج بحجم 122B على بطاقة 24GB؟ شرّحنا ATSInfer الذي يقسّم llama.cpp على مستوى الموتّرات</title><link href="https://thakicloud.github.io/ar/llmops/atsinfer-hybrid-cpu-gpu-tensor-scheduling/" rel="alternate" type="text/html" title="نموذج بحجم 122B على بطاقة 24GB؟ شرّحنا ATSInfer الذي يقسّم llama.cpp على مستوى الموتّرات" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/ar/llmops/atsinfer-hybrid-cpu-gpu-tensor-scheduling</id><content type="html" xml:base="https://thakicloud.github.io/ar/llmops/atsinfer-hybrid-cpu-gpu-tensor-scheduling/"><![CDATA[<p>هذه المقالة موجّهة للمهندسين الذين يوازنون إمكانية تشغيل نموذج كبير ذاتياً على بطاقة GPU استهلاكية واحدة، ولمسؤولي البنية التحتية الذين يقرّرون مدى تصديق تغريدات “شغّل 120B على 24GB”. بدايةً: فكرة ATSInfer الأساسية (arXiv:2607.10183)، الصادرة عن باحثين في جامعة نانجينغ، بسيطة ومقنعة. حيث كان الإنزال (offloading) السابق ينقل الأشياء كتلاً على مستوى “الطبقة” أو “الخبير”، يقسّم ATSInfer وصولاً إلى <strong>الموتّرات (tensors) فرادى</strong>. ومع ذلك، فإن الرقم البارز “حتى 3.29 ضعفاً” يقوم على عدة افتراضات، والشيفرة لم تُنشر بعد. لم نُعِد إنتاج بطاقة RTX 4090 تشغّل نموذجاً بحجم 120B هنا، لذا فإن كل رقم في هذه المقالة هو <strong>قيمة أبلغت عنها الورقة</strong>، مذكورة على هذا الأساس.</p>

<h2 id="نظرة-عامة">نظرة عامة</h2>

<p>كل من شغّل نموذج LLM محلياً يصطدم بالجدار نفسه: إذا كانت أوزان النموذج أكبر من ذاكرة الـ GPU، يفيض الزائد إلى ذاكرة الـ CPU. وراية <code class="language-plaintext highlighter-rouge">-ngl</code> في llama.cpp (عدد الطبقات التي توضع على الـ GPU) تفعل هذا بالضبط. المشكلة أنها تقطع فقط على <strong>مستوى الطبقة</strong>. تخلط الطبقة الواحدة موتّرات مختلفة الطبيعة تماماً (أوزان attention، وأوزان FFN، ومعاملات التطبيع)، ومع ذلك تُعامَل بمبدأ الكل أو لا شيء: إما أن يذهب الكل إلى الـ GPU أو يبقى على الـ CPU.</p>

<p>سبب خسارة هذا الوضع الكتلي بسيط. قد يجعل نفس الـ 1GB في الـ VRAM موتّراً أسرع 10 أضعاف وآخر أسرع ضعفين فقط. الـ VRAM مورد نادر، والقطع الكتلي يمنعك من انتقاء الموتّرات ذات الربح الأعلى لكل GB. يستهدف ATSInfer هذا بالضبط. فهو يحلّل أداء كل موتّر على الـ CPU والـ GPU ويملأ الـ VRAM بدءاً من الموتّرات ذات <strong>أعلى ربح سرعة لكل GB</strong>. إذا كان ktransformers الأخير حيلةً على “مستوى الخبير” تدفع خبراء MoE إلى الـ CPU (انظر <a href="/ar/llmops/ktransformers-moe-offload-28x-validation/">مقالتنا ذات الصلة: إعادة إنتاج الـ 28 ضعفاً لـ ktransformers</a>)، فإن ATSInfer هو التعميم الأدق على “مستوى الموتّر”. واللافت أنه ينطبق على النماذج الكثيفة (dense) وليس MoE فقط.</p>

<h2 id="ما-هذه-التقنية">ما هذه التقنية</h2>

<p>ATSInfer نظام استدلال هجين CPU-GPU مبنيّ كامتداد لـ llama.cpp بنحو 15 ألف سطر من C++. كما يقول الاسم، فإن “الجدولة الآلية للموتّرات (Automated Tensor Scheduling)” هي الجوهر، وتتشابك ثلاث آليات.</p>

<pre><code class="language-mermaid">flowchart TB
    A["أوزان النموذج&lt;br/&gt;(RAM، تتجاوز سعة VRAM)"] --&gt; B["تحليل لكل موتّر&lt;br/&gt;قياس ربح السرعة لكل GB"]
    B --&gt; C{"التوزيع الثابت&lt;br/&gt;الموتّرات الأعلى ربحاً إلى VRAM أولاً"}
    C --&gt;|"موتّرات عالية الربح"| D["مقيمة في VRAM"]
    C --&gt;|"موتّرات منخفضة الربح"| E["مقيمة في RAM"]
    D --&gt; F["نقل ديناميكي واعٍ بالحمل&lt;br/&gt;ترقية/تنزيل حسب حمل التشغيل"]
    E --&gt; F
    F --&gt; G["تنسيق غير متزامن بين CPU وGPU&lt;br/&gt;تداخل الحساب ونقل PCIe"]
    G --&gt; H["إخراج الرموز&lt;br/&gt;prefill · decode"]
</code></pre>

<p><strong>أولاً، التوزيع الثابت للموتّرات (static placement).</strong> قبل تحميل النموذج، يقيس مدى تسارع كل موتّر على الـ GPU، ثم يضع الموتّرات على الـ GPU بترتيب “أكبر سرعة مُعادة لكل GB من الـ VRAM مستخدَم”. هذا قريب من مسألة حقيبة الظهر (knapsack) ويستغل مباشرةً التباين على مستوى الموتّر الذي تجاهله التوزيع الكتلي.</p>

<p><strong>ثانياً، النقل الديناميكي الواعي بالحمل (load-aware dynamic transfer).</strong> التوزيع الثابت وحده لا يكفي. خلال الاستدلال الفعلي، يتغير الحمل لحظةً بلحظة مع حجم الدفعة وطول السياق وعدد الطلبات المتزامنة. يرقّي ATSInfer موتّراً معيناً من الـ RAM إلى الـ GPU أو ينزّله بناءً على حالة التشغيل. إذا كان التوزيع الثابت هو خط الانطلاق، فالنقل الديناميكي هو تغيير المسار أثناء القيادة.</p>

<p><strong>ثالثاً، التنسيق غير المتزامن بين CPU وGPU (asynchronous coordination).</strong> يُداخِل بين حساب الـ CPU وحساب الـ GPU ونقل الـ PCIe الذي يصل بينهما. التنفيذ الساذج يترك الـ GPU خاملاً ينتظر عمل الـ CPU أو حركة البيانات؛ وطبقة التنسيق هذه تملأ ذلك الوقت الخامل. تُبلّغ الورقة أن هذا يرفع متوسط استغلال SM (المعالج المتعدد التدفق) في الـ GPU بنحو 70%.</p>

<h2 id="النتائج-التي-تبلّغ-عنها-الورقة">النتائج التي تبلّغ عنها الورقة</h2>

<p>مرة أخرى: الأرقام أدناه <strong>قيم أبلغت عنها الورقة</strong>، لا شيء أعدنا إنتاجه. شيفرة ATSInfer لم تُنشر بعد (حتى نقاش التغريدات قال “أيها الباحثون، رجاءً شاركوا الشيفرة مع فريق llama.cpp”)، وإعادة إنتاج نموذج بحجم 120B مع RTX 4090 صعبة في البيئة المعزولة وراء هذه المقالة. لذا بدلاً من إعادة الإنتاج، نركّز على <strong>التحليل البنيوي والدلالات</strong>.</p>

<p>عنوان الورقة العريض: مقابل الأنظمة الهجينة الحالية (بما فيها إنزال llama.cpp على مستوى الطبقة)، يتحسّن الـ prefill (الإنتاجية حتى أول رمز) حتى 1.94 ضعفاً، والـ decode (الرموز المولّدة في الثانية) حتى 3.29 ضعفاً.</p>

<p><img src="/assets/images/atsinfer-hybrid-cpu-gpu-tensor-scheduling-results.png" alt="أقصى تسارع تبلّغ عنه الورقة لـ ATSInfer" /></p>

<p>الإعداد هو نظام RTX 4090 (24GB) وRTX 3060 مع 64GB من الـ RAM، والنماذج المُتحقَّق منها هي:</p>

<ul>
  <li>Llama 3.1-70B (INT4)</li>
  <li>Qwen3-Next-80B-A3B (INT4)</li>
  <li>Qwen3.5-122B-A10B (INT4)</li>
  <li>GPT-OSS-120B (MXFP4)</li>
</ul>

<p>فالادّعاء المركزي هو تشغيل نموذج بـ 122 مليار معامل (وهو MoE بعدد معاملات نشطة أقل بكثير) على بطاقة 24GB واحدة. اقرأ هنا أمرين منفصلين. أولاً، “3.29 ضعفاً” هو <strong>حدّ أقصى</strong> تحت ظروف محددة، لا متوسط عبر كل نموذج وكل دفعة. ثانياً، الربح يأتي جوهرياً لا من “إدخال ما لم يكن يدخل في الـ GPU” بل من “جعل حركة CPU-GPU الحتمية أذكى”. تبقى الفيزياء نفسها في أن عرض نطاق الـ PCIe هو عنق الزجاجة، لذا فإن مساهمة ATSInfer هي استخدام ذلك العرض دون هدر وتقليل الوقت الخامل للـ GPU.</p>

<h2 id="دلالات-لمنتجات-thakicloud">دلالات لمنتجات ThakiCloud</h2>

<p>منصة <strong>ai-platform</strong> من ThakiCloud بنية تحتية للذكاء الاصطناعي/تعلّم الآلة تخدم النماذج عبر بيئات عملاء متنوعة على Kubernetes وKueue. الجدولة على مستوى الموتّر مثل ATSInfer تتوافق مع اتجاه نراقبه عن كثب.</p>

<p>أولاً، <strong>اقتصاديات البيئات المحلية (on-premises) والسيادية.</strong> في السياقات التي لا يمكن أن تغادر فيها البيانات المبنى، كعملاء القطاع العام والمالي المحليين، يجب أن تعمل النماذج على بطاقات GPU مملوكة. إذا استطاعت بضع بطاقات استهلاكية حمل نموذج متوسط إلى كبير بدلاً من رف من ثماني بطاقات H100، تنخفض النفقات الرأسمالية الأولية بشكل كبير. ما تُظهره تجارب ATSInfer هو أن الافتراض “إن نقص الـ VRAM فعليك ببساطة شراء المزيد من الـ GPU” يمكن تخفيفه جوهرياً عبر تحسين توزيع الموتّرات. الثمن، بالطبع، هو انخفاض الإنتاجية، لذا فهو غير مناسب لأحمال حرجة الكمون. الحكم على هذه المقايضة لكل حمل هو دور طبقة الخدمة لدينا.</p>

<p>ثانياً، <strong>الاقتران بالجدولة متعددة المستأجرين.</strong> “النقل الديناميكي الواعي بالحمل” في ATSInfer هو حركة موتّرات داخل عقدة واحدة، لكن الفكرة تصمد على مستوى العنقود أيضاً. عندما يصفّ Kueue موارد الـ GPU ويوزّعها، فإن سياسة تقرّر أي طلب يُعالَج بأي دقة وأي ملف إنزال بناءً على الحمل هي مجال نفكّر فيه بالفعل. تماماً كما يعصر التحليل على مستوى الموتّر مكاسب الموارد داخل العقدة، يفعل مجدول العنقود الشيء نفسه عبر العقد.</p>

<p>ثالثاً، <strong>إعادة تعريف منحنى الكلفة-الجودة.</strong> في <a href="/ar/llmops/ktransformers-moe-offload-28x-validation/">مقالة إعادة إنتاج ktransformers</a> أظهرنا بالقياس المباشر أن أرقاماً بارزة مثل “28 ضعفاً” تقوم على افتراضات خفية. و”3.29 ضعفاً” في ATSInfer يستحق العدسة نفسها. ليس رقماً تسويقياً، بل القيمة التي تظهر على نماذج عملائنا الفعلية ودفعاتهم الفعلية واتفاقيات مستوى الخدمة الفعلية، هي ما نتحقق منه. التنافسية عند كلفة خدمة منخفضة تأتي في النهاية من تراكم هذا النوع من التحقق بالضبط.</p>

<h2 id="الحدود-والحجج-المضادة">الحدود والحجج المضادة</h2>

<p>أكبر حدّ هو <strong>عدم نشر الشيفرة.</strong> لا يمكن التحقق مما إذا كانت أرقام الورقة قابلة لإعادة الإنتاج، وما إذا كانت تصمد على أجهزة ونماذج أخرى، إلا بعد ظهور الشيفرة. امتداد بـ 15 ألف سطر من C++ هو أيضاً مهمة صيانة ودمج مع المنبع (upstream) غير هيّنة. الفرع (fork) الذي لا يُدمج أبداً يتخلّف عن تغييرات المنبع مع الوقت ويفقد قيمته.</p>

<p>ثانياً، <strong>اعتماد المكاسب على الظروف.</strong> يعتمد أثر التوزيع على مستوى الموتّر بشدة على أداء الـ CPU، وعرض نطاق الـ RAM، وجيل الـ PCIe. تفترض تجارب الورقة 64GB من الـ RAM؛ فإن نقص الـ RAM لا يترك مجالاً لإبقاء الموتّرات على الـ CPU أصلاً. على أنظمة PCIe 3.0، يرجّح أن يصبح النقل عنق الزجاجة ويقلّص الربح جوهرياً ([تقدير]، لا تفصّل الورقة مقارنةً جيلاً بجيل).</p>

<p>ثالثاً، <strong>السقف المتأصّل لتحسين الـ decode.</strong> الـ decode مهمة مقيّدة بالذاكرة (memory-bound). مهما أحسنت الجدولة، يجب الوصول إلى الأوزان خارج الـ VRAM بطريقة ما مع كل رمز، فيكون حتماً أبطأ من الإقامة الكاملة في الـ VRAM. ما يفعله ATSInfer هو “تقليل مقدار التباطؤ”، لا “إزالة التباطؤ”. وبناءً على الحجة المعاكسة: لخدمة إنتاجية تحتاج حقاً كموناً منخفضاً وإنتاجية عالية، يبقى استخدام بطاقة GPU تحمل النموذج كاملاً في الـ VRAM هو القرار الصائب. يتألق ATSInfer في مجال التطوير والتقييم والدفعات الصغيرة حيث لا تقدر على تلك البطاقة أو لا تحتاجها.</p>

<p>ومع ذلك، للاتجاه قيمة واضحة. عصر استغلال الموارد عبر البرمجيات بدلاً من إضافة العتاد يناسب منصةً مثل منصتنا بشكل خاص، منصةً تتخذ من العمل المحلي وكفاءة الكلفة والاستضافة الذاتية أسلحةً لها. حالما تُنشر الشيفرة، نخطط لإدماجها في معايير الخدمة لدينا وقياس الأرقام الحقيقية بأنفسنا.</p>

<h2 id="المصادر">المصادر</h2>

<ul>
  <li>ورقة ATSInfer: <a href="https://arxiv.org/abs/2607.10183">arXiv:2607.10183, Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices</a></li>
  <li>مقالة ذات صلة: <a href="/ar/llmops/ktransformers-moe-offload-28x-validation/">رفّ بـ 400 ألف دولار على 24GB؟ أعدنا إنتاج الـ 28 ضعفاً لـ ktransformers</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="llmops" /><category term="ATSInfer" /><category term="llama.cpp" /><category term="CPU-offloading" /><category term="GPU" /><category term="LLM-serving" /><category term="LLMOps" /><category term="quantization" /><category term="infrastructure" /><summary type="html"><![CDATA[بدلاً من إنزال طبقات أو خبراء كاملين، يوزّع ATSInfer الموتّرات فرادى بين الـ CPU والـ GPU. تُبلّغ الورقة عن تشغيل نماذج تتجاوز سعة الـ VRAM بكثير على بطاقة RTX 4090 واحدة مع رفع سرعة الـ decode حتى 3.29 ضعفاً. قرأنا arXiv:2607.10183 وفكّكنا آلية العمل ومعناها للخدمة (serving).]]></summary></entry><entry xml:lang="en"><title type="html">Inside Kimi Code CLI: How an Open Source Terminal Agent Swallows Editors via ACP</title><link href="https://thakicloud.github.io/en/agentops/kimi-code-cli-acp-open-source-agent/" rel="alternate" type="text/html" title="Inside Kimi Code CLI: How an Open Source Terminal Agent Swallows Editors via ACP" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/en/agentops/kimi-code-cli-acp-open-source-agent</id><content type="html" xml:base="https://thakicloud.github.io/en/agentops/kimi-code-cli-acp-open-source-agent/"><![CDATA[<p>Last week Moonshot AI released the open-weight model Kimi K3 and took the top spot on the coding leaderboards. But something that touches developer workflows even more directly slipped out quietly alongside it: <strong>Kimi Code CLI</strong>, an open source terminal coding agent that Moonshot released under the MIT license. LinkedIn timelines were full of posts framing it as “features Claude Code doesn’t have.” We didn’t take that line at face value and instead checked the official repository and documentation ourselves. The short version: half of the pitch holds up, and half is overstated. The genuinely interesting part turned out to be something the marketing didn’t emphasize at all.</p>

<p>This post covers what Kimi Code CLI is, what it actually delivers, and why it is worth watching from the perspective of a team running a Kubernetes based AI platform. We spend a good part of it on why Agent Client Protocol, an open standard, is a piece that could reshape the agent ecosystem.</p>

<h2 id="what-kimi-code-cli-is">What Kimi Code CLI Is</h2>

<p>Kimi Code CLI is an agentic coding tool that runs in the terminal, in the same family as Claude Code, Gemini CLI, and Codex CLI. The official repository is <a href="https://github.com/MoonshotAI/kimi-code">MoonshotAI/kimi-code</a>, and it evolved from the earlier project <a href="https://github.com/MoonshotAI/kimi-cli">MoonshotAI/kimi-cli</a>, carrying existing sessions and configuration forward. Both repositories are official Moonshot projects, so be careful not to confuse them with similarly named third-party projects.</p>

<p>Here is the first correction worth making. The tool’s official name is not “the CLI for Kimi K3” but <strong>Kimi Code CLI</strong>. It is not tied to a single model. By default it pairs with Moonshot’s coding-specialized model, Kimi K2.7 Code, but configuration lets you switch to K3 or other models as well. K3 is one of several models the CLI can attach to, not a model the CLI was purpose-built for. K3 itself is a 2.8 trillion parameter open MoE model that Moonshot released on July 16, 2026, built around Kimi Delta Attention and up to a 1 million token context window. This launch was covered jointly by CNBC, Bloomberg, and Forbes, among other major outlets.</p>

<p>Establishing the full picture first makes the details land much better later. The core idea is that Kimi Code CLI plays a dual role: on one side it is an MCP client connecting to tools and data, and on the other side it is an ACP server connecting to editors.</p>

<pre><code class="language-mermaid">flowchart TB
    subgraph EDITOR["Developer Editor (ACP Client)"]
        ZED["Zed"]
        JB["JetBrains family"]
        VSC["VS Code / Neovim"]
    end
    ACP["Agent Client Protocol&lt;br/&gt;JSON-RPC over stdio"]
    subgraph CLI["Kimi Code CLI (Agent Core)"]
        MAIN["Main Agent&lt;br/&gt;Maintains conversation history"]
        SUB["Sub-agents&lt;br/&gt;coder · explore · plan&lt;br/&gt;Each with an isolated context"]
    end
    MODEL["Model Layer&lt;br/&gt;Kimi K2.7 Code / K3&lt;br/&gt;or OpenAI-compatible endpoint"]
    subgraph MCP["MCP Servers (Tools · Data)"]
        T1["Context7"]
        T2["Chrome DevTools"]
        T3["Internal connectors"]
    end

    EDITOR --&gt; ACP
    ACP --&gt;|kimi acp| MAIN
    MAIN --&gt; SUB
    MAIN --&gt;|inference request| MODEL
    SUB --&gt;|inference request| MODEL
    MAIN --&gt;|tool call| MCP
</code></pre>

<h2 id="sub-agents-splitting-context-to-keep-the-main-thread-clean">Sub-agents: Splitting Context to Keep the Main Thread Clean</h2>

<p>Kimi Code CLI ships with three built-in sub-agents. <strong>coder</strong> is the general engineering agent that reads and writes files and runs commands to make actual changes. <strong>explore</strong> is read-only and dedicated to surveying the codebase. <strong>plan</strong> produces implementation plans and architectural designs without running any shell commands. This split is spelled out in the official <a href="https://moonshotai.github.io/kimi-code/en/customization/agents.html">Agents and Sub-Agents</a> documentation.</p>

<p>What matters here is not the naming but the context isolation. Each sub-agent gets a fully independent context window and sees only the task description the main agent explicitly hands it. The main agent’s conversation history is never exposed to a sub-agent, and the intermediate reasoning and tool call logs a sub-agent produces never leak back into the main history. A sub-agent returns only its final conclusion. That is why the main context stays thin instead of bloating with logs over a long session. Background and parallel execution are also supported, so multiple exploration tasks can run at once and their results flow back automatically once complete.</p>

<p>This pattern is not unfamiliar to us. The internal orchestration harness behind this blog also delegates exploration to low-cost sub-agents and pulls back only summaries to protect the main context. The principle that context hygiene is both a cost and a quality lever holds regardless of which tool implements it.</p>

<h2 id="mcp-a-configuration-experience-without-hand-editing-json">MCP: A Configuration Experience Without Hand-Editing JSON</h2>

<p>Model Context Protocol integration is managed through two paths. The first is the CLI subcommand set: <code class="language-plaintext highlighter-rouge">kimi mcp add</code>, <code class="language-plaintext highlighter-rouge">kimi mcp list</code>, <code class="language-plaintext highlighter-rouge">kimi mcp remove</code>, and <code class="language-plaintext highlighter-rouge">kimi mcp authorize</code> manage servers. For example, you can attach a documentation search server over HTTP transport, or a browser automation server over stdio transport.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># HTTP transport (with optional OAuth)</span>
kimi mcp add <span class="nt">--transport</span> http context7 https://mcp.context7.com/mcp

<span class="c"># stdio transport connecting to a local process</span>
kimi mcp add <span class="nt">--transport</span> stdio chrome-devtools <span class="nt">--</span> npx chrome-devtools-mcp@latest
</code></pre></div></div>

<p>The second path is the interactive slash command <code class="language-plaintext highlighter-rouge">/mcp-config</code> inside the TUI, which lets you add, edit, and authenticate servers without touching a JSON config file directly. <code class="language-plaintext highlighter-rouge">/mcp</code> shows the currently connected servers and the list of loaded tools. The claim that “you don’t need to edit JSON directly,” which the LinkedIn post emphasized, is accurate. That said, this convenience by itself is not something Claude Code lacks either; we return to that point later in this post. See the <a href="https://moonshotai.github.io/kimi-cli/en/customization/mcp.html">MCP configuration</a> docs for details.</p>

<h2 id="agent-client-protocol-the-most-important-piece-of-this-tool">Agent Client Protocol: The Most Important Piece of This Tool</h2>

<p>This is the most interesting part of this article. Agent Client Protocol, abbreviated ACP, is an open standard built by the Zed editor team. It is Apache licensed and runs JSON-RPC 2.0 over stdio, with the editor spawning the agent as a child process and communicating through standard input and output. The transport mechanism itself is identical to the Language Server Protocol.</p>

<p>An analogy helps a great deal here. Before LSP existed, every editor needed a separate integration for every language. LSP turned that M-times-N problem into M-plus-N: an editor only needs to implement the standard once, and it benefits from every language server anyone builds. ACP does exactly the same thing for agents. Once an editor implements ACP, any agent built by anyone plugs into it through a standard interface. This concept is explained in <a href="https://zed.dev/acp">Zed’s introduction to ACP</a> and in <a href="https://blog.marcnuri.com/agent-client-protocol-acp-introduction">Marc Nuri’s write-up</a>.</p>

<p>It is easy to confuse ACP with MCP, but the direction is reversed. MCP flows from the agent toward tools and data, with the agent acting as the MCP client. ACP flows from the editor toward the agent, with the agent acting as the ACP server and the editor acting as the ACP client. The same agent can play the MCP client role on one side and the ACP server role on the other at the same time, which is exactly why the diagram above draws that dual role.</p>

<p>Kimi Code CLI supports this protocol natively through the <code class="language-plaintext highlighter-rouge">kimi acp</code> subcommand, with no separate installation required. Zed connects natively, JetBrains connects through a plugin, and based on Zed’s ACP registry, several other editor integrations are already listed. Developers can drive a Kimi session without ever leaving the editor they already use.</p>

<h2 id="image-and-video-input-what-actually-works">Image and Video Input: What Actually Works</h2>

<p>The LinkedIn post claimed you can “feed a screen capture directly as input.” This needs a correction. The feature Moonshot actually markets front and center is not a static screenshot but <strong>screen recording video input</strong>. The repository description says that dropping a screen recording or demo clip into the chat lets the agent directly see and understand behavior that’s hard to describe in words. Pasting images into the CLI input is also supported, and the default model, Kimi K2.7 Code, is natively multimodal with a 400 million parameter vision encoder called MoonViT, so it accepts text, images, and video all together. That said, when you attach a custom model, that model’s modalities must explicitly declare image support for this to work correctly. To summarize, image input does work, but the feature actually marketed as the real differentiator is video input, and the word “screenshot” is somewhat inaccurate here.</p>

<h2 id="installation-is-really-three-steps">Installation Is Really Three Steps</h2>

<p>The installation flow is as simple as advertised. The commands below follow the <a href="https://moonshotai.github.io/kimi-cli/en/guides/getting-started.html">official getting started guide</a>. Our internal sandbox does not have access to the relevant distribution domain, so we did not run these ourselves and are not logging execution output. We are not fabricating any benchmark numbers here; we are only citing verified commands.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># 1) Run the install script (also installs uv)</span>
curl <span class="nt">-LsSf</span> https://code.kimi.com/install.sh | bash

<span class="c"># 2) Run it from your project directory</span>
kimi

<span class="c"># 3) Set up authentication</span>
/login
</code></pre></div></div>

<p>macOS also supports <code class="language-plaintext highlighter-rouge">brew install kimi-code</code>, and Windows has a PowerShell script. Building from source requires Node 24.15 or later and pnpm. The license is MIT, which puts few restrictions on reading the code, forking it, or deploying it internally.</p>

<h2 id="model-and-provider-openness">Model and Provider Openness</h2>

<p>Context length tops out at 256,000 tokens for the K2.6 line, and Moonshot’s marketing claims up to 1 million tokens for K3. More important is provider openness. In <code class="language-plaintext highlighter-rouge">~/.kimi-code/config.toml</code> you can register multiple providers, including OpenAI-compatible endpoints, an Anthropic API key, and Google GenAI or Vertex AI. That means the CLI is not locked into a single model. It also automatically handles the <code class="language-plaintext highlighter-rouge">reasoning_content</code> field from third-party reasoning models. See <a href="https://moonshotai.github.io/kimi-cli/en/configuration/providers.html">Providers and models</a> for the details.</p>

<h2 id="is-this-a-feature-claude-code-lacks-an-honest-comparison">Is This a Feature Claude Code Lacks? An Honest Comparison</h2>

<p>The most widely circulated pitch was “this gives you features Claude Code doesn’t have.” After verifying it, that framing turns out to be mostly overstated.</p>

<p>Sub-agents and context isolation are already offered by Claude Code through its own sub-agent feature, in the same way. MCP is also mature in Claude Code, which already supports stdio, SSE, and HTTP transports. Image pasting exists in Claude Code as well. None of these three are actual differentiators.</p>

<p>The real differences sit in two places. First, how ACP support is delivered. Kimi Code CLI bakes ACP into the CLI itself as a first-class feature through the <code class="language-plaintext highlighter-rouge">kimi acp</code> subcommand. Claude Code, by contrast, connects through a separate adapter package built by Zed, currently in beta. From a user’s perspective, the former just works once you turn it on, while the latter requires bolting on an additional bridge. Second, model openness. Kimi is open across its open-weight K series with the ability to switch providers, while Claude Code is limited to Anthropic’s own models. A third difference follows from this: self-hosting potential. Kimi combines an open source CLI with open-weight models, making on-premise serving possible, whereas Claude Code has an open CLI but a model that is API-only. The relevant evidence is in <a href="https://zed.dev/blog/claude-code-via-acp">Zed’s post on the Claude Code ACP beta</a>.</p>

<h2 id="implications-for-thakiclouds-products">Implications for ThakiCloud’s Products</h2>

<p>This topic touches both the agent tooling axis and the open model and on-premise infrastructure axis, so we apply two lenses together.</p>

<p>Through the Paxis lens, Kimi Code CLI’s structure overlaps quite a bit with our product’s design direction. Paxis is ThakiCloud’s Agent-Native Cloud control plane, treating skills, tools, policies, and audit logs as first-class resources. The way Kimi’s coder, explore, and plan sub-agents run in parallel with isolated contexts shares the same underlying philosophy as the way Paxis’s skill harness selects among more than 960 skills with BM25 and runs them in isolated sandboxes. ACP in particular, as a vendor-neutral standard, is a direct opportunity for Paxis. Any agent we deploy, including ones running our own fine-tuned models, could plug into a customer’s development editor such as Zed or JetBrains through a standard interface simply by implementing ACP. The combination of MCP for connecting to data and ACP for connecting to editors is exactly the kind of integrated picture we are aiming for.</p>

<p>Through the ai-platform lens, openness translates directly into deployment freedom. Running an open-weight K-series model on top of our cluster’s Kueue GPU scheduling and vLLM serving, and routing the CLI to an internal endpoint, would let us build an internal coding agent without depending on an external API or exporting data outside our environment. That aligns well with on-premise security requirements in domains like finance or the public sector where code cannot leave the organization, including requirements tied to Korea’s National Intelligence Service. As capability becomes commoditized and cheaper, what enterprises actually end up paying for is a controlled execution environment, a point we have made in earlier posts as well. Kimi Code CLI matters precisely because it opens up that execution layer as open source.</p>

<h2 id="caveats-and-counterarguments">Caveats and Counterarguments</h2>

<p>A few things deserve a colder look. First, some third-party deep dives reference internal engine names or layer structures that don’t appear in the official documentation, which suggests they may be the product of reverse engineering. It is safer to treat the official docs as the source of truth before citing those as fact. Second, there are community reports that the ACP path produces better response quality than other connection methods, but this is anecdotal rather than benchmarked, and we don’t treat it as verified data. Third, even with open weights, actually serving a 2.8 trillion parameter class model on-premise requires substantial GPU resources. Openness does not automatically mean easy self-hosting, and the API route remains the practical choice for small teams. Fourth, the maturity and stability of the tooling ecosystem may still favor Claude Code or Codex CLI. Being open source does not by itself mean production ready.</p>

<p>Even so, the direction toward agents and editors loosely coupled through open standards is a clear trend. A world where developers aren’t locked into a single vendor’s CLI, and can swap out the model and the editor independently, is a better world for developers. Kimi Code CLI is one of the pieces bringing that world closer.</p>

<h2 id="sources">Sources</h2>

<ul>
  <li><a href="https://github.com/MoonshotAI/kimi-code">MoonshotAI/kimi-code (official repository)</a></li>
  <li><a href="https://github.com/MoonshotAI/kimi-cli">MoonshotAI/kimi-cli (previous repository)</a></li>
  <li><a href="https://moonshotai.github.io/kimi-cli/en/guides/getting-started.html">Kimi Code CLI getting started guide</a></li>
  <li><a href="https://moonshotai.github.io/kimi-code/en/customization/agents.html">Agents and Sub-Agents documentation</a></li>
  <li><a href="https://moonshotai.github.io/kimi-cli/en/customization/mcp.html">MCP configuration documentation</a></li>
  <li><a href="https://zed.dev/acp">Zed - Agent Client Protocol</a></li>
  <li><a href="https://blog.marcnuri.com/agent-client-protocol-acp-introduction">ACP: The LSP for AI Coding Agents</a></li>
  <li><a href="https://zed.dev/blog/claude-code-via-acp">Zed - Claude Code via ACP (beta)</a></li>
  <li><a href="https://www.marktechpost.com/2026/07/16/moonshot-ai-releases-kimi-k3-a-2-8-trillion-parameter-open-moe-model-with-kimi-delta-attention-and-1m-context/">MarkTechPost - Kimi K3 launch</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="agentops" /><category term="agentops" /><category term="kimi" /><category term="moonshot" /><category term="coding-agent" /><category term="mcp" /><category term="agent-client-protocol" /><category term="paxis" /><category term="thakicloud" /><summary type="html"><![CDATA[We dig into the open source coding CLI that Moonshot AI released alongside Kimi K3, working strictly from the official docs and repository. From coder/explore/plan sub-agents to interactive MCP configuration and its real differentiator, native Agent Client Protocol support, we verify how much of the 'features Claude Code lacks' marketing line actually holds up.]]></summary></entry><entry xml:lang="en"><title type="html">The Question Behind a 25 Billion Won Bill: Why AI Agent Costs Have Become Invisible</title><link href="https://thakicloud.github.io/en/llmops/agent-cost-observability-billing-crisis/" rel="alternate" type="text/html" title="The Question Behind a 25 Billion Won Bill: Why AI Agent Costs Have Become Invisible" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/en/llmops/agent-cost-observability-billing-crisis</id><content type="html" xml:base="https://thakicloud.github.io/en/llmops/agent-cost-observability-billing-crisis/"><![CDATA[<p>This piece is written for platform and infrastructure engineers rolling out Claude Code or AI agents across an organization, and for finance and procurement owners who have to explain next month’s AI bill. The short version: the flood of AI cost news over the past month is not really a story about “AI being expensive.” The real problem is that the bill doesn’t explain what it’s actually for. In an agent architecture, a single user request can fan out into dozens or hundreds of model calls, tool executions, and automatic retries on failure, and the final dollar amount alone gives you no way to tell which loop leaked the money. We think this observability gap is the real source of the pain the market is feeling right now.</p>

<h2 id="overview">Overview</h2>

<p>If you had to sum up the news from late June through July 2026 in one line, it would be this: using frontier models indiscriminately for every task is unsustainably expensive, and it’s hard to even trace where that cost is coming from. The clearest illustration was dramatic. A user in Korea first saw a charge attempt of roughly 2.5 billion won, then one for roughly 25 billion won. No money actually left the user’s account, since the charge exceeded the card’s limit, but the fact that an abnormal figure was pushed all the way to the card authorization stage, not just displayed as a UI glitch, is what gives the incident its weight.</p>

<p>Around the same time, news broke at other levels of the stack. An AI cost auditing firm reviewed the bills of 60 companies and alleged significant overcharging, several large enterprises began splitting frontier-model usage across cheaper models depending on the task, and reports emerged that companies in the US and Europe were migrating to Chinese open-weight models to cut costs. Interestingly, model vendors themselves began responding in a way that essentially conceded the point, that running the top-tier model on everything, all the time, isn’t sustainable. Stories from several different directions were converging on the same spot.</p>

<h2 id="what-happened-over-the-past-month">What Happened Over the Past Month</h2>

<p>The first issue to surface was the reliability of the billing system itself. According to ZDNet Korea, a charge to a Korean university student started at roughly $1.66 million and then jumped tenfold to roughly $16.62 million. Anthropic later explained that the auto-recharge amount had been set to an abnormally high figure by mistake, but the user maintained that he had never configured auto-recharge in the first place, so why that setting existed remains unexplained. The detail that the user sent more than fifteen emails to multiple departments and only received an automated reply four days later says more about the gap in the response system than about the technical error itself.</p>

<p>The second issue was the observability of agent costs. According to AI Times and The Information, cost-auditing startup Vaudit reviewed roughly $34 million in bills across 60 companies and claimed that about $1.7 million of it was overcharged. A large share of what it reviewed was Claude Code usage, and companies including Panasonic, HP, and Honda were named as clients. The patterns Vaudit pointed to included calls made on a cheap model but billed at a premium model’s rate, charges for jobs that never actually completed, and repeated automatic retries after failures, what it called a retry storm. Two caveats are worth stating clearly here. First, Anthropic pushed back, saying it does not bill for incomplete requests or error responses and that it has seen no evidence of widespread overcharging. Second, Vaudit is a commercial auditing firm that takes a cut of any refunds it wins, so these figures are best read as one party’s findings rather than an independent audit. In other words, this is currently a standoff between an auditor’s allegations and a vendor’s denial.</p>

<p>The third was how the market responded. The Information reported that companies had started splitting workloads by task, routing simple classification, summarization, and transformation jobs to cheaper models, complex coding and agentic work to frontier models, and repetitive high-volume jobs to open-weight or self-hosted models. The Financial Times reported that DoorDash, Siemens, and Airbnb had adopted DeepSeek or Moonshot-family models to cut costs. Business Insider reported that even Anthropic’s own platform executives acknowledged that shadow IT, business units adopting AI tools on their own without coordination, had driven runaway AI spending at some companies, though their prescribed fix was task-level model selection and centralized cost governance rather than usage bans or blanket budget caps. Pricing policy itself kept shifting too. Whether the latest high-performance models were included in subscriptions, and when usage-based billing would kick in, was adjusted repeatedly, and promotional deadlines kept getting extended. One example: the free tier for Claude Fable 5 was reportedly extended through July 19. For procurement teams, the bigger headache wasn’t performance, it was not being able to predict next month’s bill.</p>

<h2 id="why-agent-costs-go-unobserved">Why Agent Costs Go Unobserved</h2>

<p>There’s a single thread running through all three strands of news. In agent workloads, the distance between what the user sees and what the bill records has grown too large. A traditional API call was one request, one response, one line of cost. A coding agent or agent SDK, by contrast, expands a single instruction into planning, tool calls, file edits, verification, and automatic retries on failure. That expansion happens somewhere the user never sees, and the bill records only the sum of it, in one line.</p>

<pre><code class="language-mermaid">flowchart TB
    U["1 user request"] --&gt; P["Agent planning"]
    P --&gt; L["Execution loop"]
    L --&gt; T["Tool calls · model calls&lt;br/&gt;tens to hundreds of times"]
    T --&gt; R{"Success?"}
    R --&gt;|"Failure"| RS["Automatic retry&lt;br/&gt;(retry storm)"]
    RS --&gt; T
    R --&gt;|"Success"| ACC["Tokens · cache · tool calls&lt;br/&gt;aggregated"]
    ACC --&gt; INV["Bill: one final line item"]
    INV -.observability gap.-&gt; U
</code></pre>

<p>In this structure, most of the points where cost leaks out sit outside the user’s field of view. A retry loop quietly runs and inflates the call count, an intermediary cloud provider sits between the actual model usage and the final billing record and creates drift between the two, and one misconfigured setting, like auto-recharge, can push an abnormal amount all the way to the card authorization stage. The three stories may look unrelated, but they’re all different faces of the same observability gap. That’s why setting a monthly cap per user isn’t enough on its own. What’s actually needed is an instrumentation layer that captures per-model cost, per-session tokens, cache tokens, tool call counts, failure and retry costs, and day-over-day anomaly rates centrally, at the moment each call happens. Without observability there’s no control, and without control the bill will always be the document that surprises you after the fact.</p>

<h2 id="implications-for-thakiclouds-products">Implications for ThakiCloud’s Products</h2>

<p>This is a problem that two of ThakiCloud’s products each target from a different angle. Because the infrastructure lens and the agent lens complement each other here, we apply both to this topic.</p>

<p><strong>The ai-platform lens: owning repetitive workloads is the answer.</strong> The market has arrived at a clear conclusion. Running even trivial tasks on frontier models is unsustainably expensive, and self-hosting an open-weight model is the economical choice for repetitive, high-volume work. ThakiCloud’s ai-platform is Kubernetes-based AI/ML infrastructure built for exactly that gap. It queues GPUs with Kueue to push utilization higher, serves open-weight models through vLLM, and separates usage and billing by department through multi-tenant isolation. Where usage-based API pricing produces unpredictable bills, self-hosting builds a structure on top of a fixed GPU cost where the unit price doesn’t spike even as usage grows. Unlike external usage-based pricing that keeps shifting, on-premises and sovereign deployments turn cost predictability itself into an asset. And the fact that data never leaves the organization is an additional value for organizations facing heavy domestic regulatory and security requirements.</p>

<p><strong>The Paxis lens: making every agent action auditable.</strong> The core of the observability gap was the agent loop, and that is precisely the territory Paxis addresses. Paxis is ThakiCloud’s Agent-Native Cloud control plane, running on top of ai-platform, that treats Skills, Tools, Policies, and Audit Logs as first-class resources. Which skill an agent invoked, through which tool, how many times, and in which sandbox it ran, all of it is captured in an audit log. In this structure, a retry storm can’t quietly inflate a bill: the retry loop shows up directly in the audit trail, and policy gates block calls that cross a threshold. A design that selects from more than 960 skills using BM25, runs them in isolated sandboxes, and routes every action through policy and audit is a structural answer to exactly the problem of not being able to tell, from the bill alone, which loop generated the cost. Low-cost serving through ai-platform makes agents economical, and action-level observability through Paxis makes that economics predictable. That’s how the two lenses fit together.</p>

<h2 id="limitations-and-counterarguments">Limitations and Counterarguments</h2>

<p>In the interest of balance, let’s state the counterargument clearly. First, frontier models aren’t automatically wasteful. According to the Wall Street Journal, companies like Shopify believe that for complex coding and multi-step agent work, a frontier model’s higher price can be justified if it saves enough engineering time. Spotify and Twilio, by contrast, are weighing more carefully whether a marginal performance gain justifies the added cost. The takeaway isn’t “abandon frontier models,” it’s “split the workload by task difficulty.” Self-hosting isn’t a universal answer either. Pushing tasks that require the highest level of reasoning down to open-weight models degrades quality, and it introduces a new operational burden of GPU operations, model updates, and security patching.</p>

<p>Second, the overbilling figures cited in this piece are not settled facts. Vaudit’s claims come from a commercial auditing firm, and Anthropic has denied them, so the accurate reading right now is that the two sides’ positions are in direct conflict. In the 25-billion-won billing incident too, no money actually changed hands, and no technical explanation for why the auto-recharge setting existed has been made public. The conclusion we draw from this news isn’t aimed at any particular vendor. It’s a principle: in the agent era, whichever vendor you use, cost observability and governance need to be secured on the user’s side. Choosing a good model and governing that model are two separate problems, and the news of the past month simply exposed the fact that the latter has been an empty seat all along.</p>

<h2 id="sources">Sources</h2>

<ul>
  <li><a href="https://zdnet.co.kr/view/?no=20260709165452">ZDNet Korea, “Korean User Hit With 25 Billion Won Payment Request, Anthropic Billing Error Controversy” (2026-07-09)</a></li>
  <li><a href="https://zdnet.co.kr/view/?no=20260716093004">ZDNet Korea, “Anthropic’s 25 Billion Won Charge Turns Out to Be an Auto-Recharge Configuration Error” (2026-07-16)</a></li>
  <li><a href="https://www.aitimes.com/news/articleView.html?idxno=212155">AI Times, “Anthropic Faces ‘AI Overbilling’ Controversy, Charged for Failed Jobs Too” (2026-06)</a></li>
  <li><a href="https://www.theinformation.com/titv/fedld">The Information, report on enterprises adopting AI cost controls and model diversification (2026-06-23)</a></li>
  <li><a href="https://www.ft.com/content/9c8ff45b-7c20-4c2e-93c9-c52339ffdcee">Financial Times, “Companies turn to Chinese AI models to cut costs” (2026-07)</a></li>
  <li><a href="https://www.businessinsider.com/anthropic-ai-costs-responses-routers-2026-7">Business Insider, “Anthropic Official Warns Against ‘Wrong’ AI Cost Response” (2026-07-15)</a></li>
  <li><a href="https://www.wsj.com/cio-journal/meet-the-companies-shelling-out-for-top-ai-models-e1fe3375">The Wall Street Journal, “Meet the Companies Shelling Out for Top AI Models” (2026-07)</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="llmops" /><category term="LLMOps" /><category term="FinOps" /><category term="AgentCost" /><category term="CostObservability" /><category term="ModelRouting" /><category term="self-hosting" /><category term="Paxis" /><category term="AIInfrastructure" /><summary type="html"><![CDATA[From a 25 billion won bill sent to a user in Korea to allegations of overbilling at 60 companies, the past month of AI cost news points to a single blind spot. In the age of agents, where a single request can trigger hundreds of model calls, the bill no longer explains what you actually paid for. Here is how we think that observability gap should be closed.]]></summary></entry><entry xml:lang="en"><title type="html">A 122B Model on a 24GB Card? We Dissected ATSInfer, Which Slices llama.cpp at Tensor Granularity</title><link href="https://thakicloud.github.io/en/llmops/atsinfer-hybrid-cpu-gpu-tensor-scheduling/" rel="alternate" type="text/html" title="A 122B Model on a 24GB Card? We Dissected ATSInfer, Which Slices llama.cpp at Tensor Granularity" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/en/llmops/atsinfer-hybrid-cpu-gpu-tensor-scheduling</id><content type="html" xml:base="https://thakicloud.github.io/en/llmops/atsinfer-hybrid-cpu-gpu-tensor-scheduling/"><![CDATA[<p>This post is for engineers weighing whether to self-serve a large model on a single consumer GPU, and for infra owners deciding how much to trust the “run 120B on 24GB” tweets going around. Up front: the core idea of ATSInfer (arXiv:2607.10183), released by researchers at Nanjing University, is simple and persuasive. Where prior offloading moved things in chunks at the granularity of a “layer” or an “expert,” ATSInfer slices down to <strong>individual tensors</strong>. That said, the headline “up to 3.29x” rests on a few premises, and the code is not yet public. We did not reproduce an RTX 4090 running a 120B-class model here, so every number in this post is a <strong>value reported by the paper</strong>, stated as such.</p>

<h2 id="overview">Overview</h2>

<p>Anyone who has run a local LLM hits the same wall: if the model weights are larger than GPU memory, the overflow spills to CPU memory. llama.cpp’s <code class="language-plaintext highlighter-rouge">-ngl</code> flag (how many layers to place on the GPU) does exactly this. The problem is that it cuts only at <strong>layer granularity</strong>. A single layer mixes tensors of very different character (attention weights, FFN weights, normalization params), yet they are handled all-or-nothing: either the whole thing goes to the GPU or it stays on the CPU.</p>

<p>Why this chunked placement loses is simple. The same 1GB in VRAM might make one tensor 10x faster and another only 2x faster. VRAM is a scarce resource, and chunked cuts prevent you from picking the tensors with the highest gain per GB. ATSInfer targets exactly this. It profiles each tensor’s CPU and GPU performance and fills VRAM starting from the tensors with the <strong>highest speed gain per GB</strong>. If the recent ktransformers was an “expert-level” trick that pushes MoE experts to the CPU (see our <a href="/en/llmops/ktransformers-moe-offload-28x-validation/">related post: reproducing the ktransformers 28x</a>), ATSInfer is the finer, “tensor-level” generalization. Notably, it applies to dense models, not just MoE.</p>

<h2 id="what-is-this-technology">What is this technology</h2>

<p>ATSInfer is a hybrid CPU-GPU inference system built as an extension to llama.cpp in roughly 15,000 lines of C++. As the name says, “Automated Tensor Scheduling” is the core, and three mechanisms interlock.</p>

<pre><code class="language-mermaid">flowchart TB
    A["Model weights&lt;br/&gt;(RAM, exceed VRAM capacity)"] --&gt; B["Per-tensor profiling&lt;br/&gt;measure speed gain per GB"]
    B --&gt; C{"Static placement&lt;br/&gt;highest-gain tensors to VRAM first"}
    C --&gt;|"high-gain tensors"| D["Resident in VRAM"]
    C --&gt;|"low-gain tensors"| E["Resident in RAM"]
    D --&gt; F["Load-aware dynamic transfer&lt;br/&gt;promote/demote by runtime load"]
    E --&gt; F
    F --&gt; G["Asynchronous CPU-GPU coordination&lt;br/&gt;overlap compute and PCIe transfer"]
    G --&gt; H["Token output&lt;br/&gt;prefill · decode"]
</code></pre>

<p><strong>First, static tensor placement.</strong> Before loading the model, it benchmarks how much each tensor speeds up on the GPU, then places tensors on the GPU in order of “most speed returned per GB of VRAM used.” This is close to a knapsack optimization and directly exploits the tensor-level heterogeneity that chunked placement ignored.</p>

<p><strong>Second, load-aware dynamic transfer.</strong> Static placement alone is not enough. During real inference, load shifts moment to moment with batch size, context length, and concurrency. ATSInfer promotes a given tensor from RAM to GPU, or demotes it, based on the runtime situation. If static placement is the starting line, dynamic transfer is changing lanes while driving.</p>

<p><strong>Third, asynchronous CPU-GPU coordination.</strong> It overlaps CPU compute, GPU compute, and the PCIe transfer that connects them. A naive implementation leaves the GPU idling while it waits on CPU work or data movement; this coordination layer fills that idle time. The paper reports this raises average GPU SM (streaming multiprocessor) utilization by about 70%.</p>

<h2 id="results-the-paper-reports">Results the paper reports</h2>

<p>Again: the numbers below are <strong>values reported by the paper</strong>, not something we reproduced. ATSInfer’s code is not yet public (even the tweet chatter said “researchers, please share the code with the llama.cpp team”), and a 120B-class model plus an RTX 4090 is hard to reproduce in the sandbox behind this post. So instead of reproduction, we focus on <strong>structural analysis and implications</strong>.</p>

<p>The paper’s headline: versus existing hybrid systems (including llama.cpp’s layer-level offloading), prefill (throughput to first token) improves by up to 1.94x, and decode (tokens generated per second) by up to 3.29x.</p>

<p><img src="/assets/images/atsinfer-hybrid-cpu-gpu-tensor-scheduling-results.png" alt="Max speedup ATSInfer reports in the paper" /></p>

<p>The setup is an RTX 4090 (24GB) and RTX 3060 system with 64GB RAM, and the validated models are:</p>

<ul>
  <li>Llama 3.1-70B (INT4)</li>
  <li>Qwen3-Next-80B-A3B (INT4)</li>
  <li>Qwen3.5-122B-A10B (INT4)</li>
  <li>GPT-OSS-120B (MXFP4)</li>
</ul>

<p>So the central claim is running a 122B-parameter model (an MoE with far fewer active parameters) on a single 24GB card. Read two things separately here. First, “3.29x” is a <strong>maximum</strong> under specific conditions, not an average across every model and batch. Second, the gain fundamentally comes not from “fitting what didn’t fit into the GPU” but from “making the unavoidable CPU-GPU traffic smarter.” The physics that PCIe bandwidth is the bottleneck stays the same, so ATSInfer’s contribution is using that bandwidth without waste and reducing GPU idle time.</p>

<h2 id="implications-for-thakicloud-products">Implications for ThakiCloud products</h2>

<p>ThakiCloud’s <strong>ai-platform</strong> is an AI/ML infrastructure that serves models across diverse customer environments on Kubernetes and Kueue. Tensor-level scheduling like ATSInfer aligns with a trend we watch closely.</p>

<p>First, <strong>the economics of on-premises and sovereign environments.</strong> In settings where data cannot leave the premises, such as domestic public-sector and financial customers, models must run on owned GPUs. If a few consumer GPUs can carry a mid-to-large model instead of a rack of eight H100s, initial CAPEX drops dramatically. What ATSInfer’s experiments show is that the premise “if VRAM is short you must simply buy more GPUs” can be substantially relaxed through tensor placement optimization. The cost, of course, is reduced throughput, so it is unsuitable for latency-critical workloads. Judging that trade-off per workload is the job of our serving layer.</p>

<p>Second, <strong>coupling with multi-tenant scheduling.</strong> ATSInfer’s “load-aware dynamic transfer” is tensor movement within a single node, but the idea holds at cluster scale too. When Kueue queues and allocates GPU resources, a policy that decides which request to handle at which precision and offloading profile based on load is an area we already think about. Just as tensor-level profiling squeezes resource gains within a node, the cluster scheduler does the same across nodes.</p>

<p>Third, <strong>redefining the cost-quality curve.</strong> In our <a href="/en/llmops/ktransformers-moe-offload-28x-validation/">ktransformers reproduction post</a> we showed by direct measurement that headline figures like “28x” rest on hidden premises. ATSInfer’s “3.29x” deserves the same lens. Not a marketing number, but the value that emerges on our customers’ actual models, actual batches, and actual SLAs, is what we verify. Competitiveness at low serving cost ultimately comes from accumulating exactly this kind of verification.</p>

<h2 id="limits-and-counterarguments">Limits and counterarguments</h2>

<p>The biggest limit is <strong>the code being unreleased.</strong> Whether the paper’s numbers reproduce, and whether they hold on other hardware and models, can only be checked once code appears. A 15,000-line C++ extension is also a nontrivial maintenance and upstream-merge task. A fork that never merges falls behind upstream changes over time and loses value.</p>

<p>Second, <strong>the condition-dependence of the gains.</strong> The effect of tensor-level placement depends heavily on CPU performance, RAM bandwidth, and PCIe generation. The paper’s experiments assume 64GB RAM; if RAM is short, there is no room to keep tensors on the CPU at all. On PCIe 3.0 systems, transfer likely becomes the bottleneck and shrinks the gain substantially ([estimate], the paper does not spell out a generation-by-generation comparison).</p>

<p>Third, <strong>the inherent ceiling on decode optimization.</strong> Decode is a memory-bound task. However well you schedule, weights outside VRAM must be accessed somehow every token, so it is inevitably slower than pure VRAM residency. What ATSInfer does is “minimize how much slower,” not “eliminate the slowdown.” Building the opposing case: for production serving that truly needs low latency and high throughput, using a GPU that holds the whole model in VRAM is still the right call. ATSInfer shines in the development, evaluation, and small-batch regime where you cannot afford that GPU, or do not need it.</p>

<p>Even so, the direction has clear value. Squeezing resource utilization through software rather than adding hardware fits a platform like ours particularly well, one that treats on-prem, cost-efficiency, and self-hosting as its weapons. Once the code is public, we plan to fold it into our serving benchmarks and measure the real numbers ourselves.</p>

<h2 id="sources">Sources</h2>

<ul>
  <li>ATSInfer paper: <a href="https://arxiv.org/abs/2607.10183">arXiv:2607.10183, Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices</a></li>
  <li>Related post: <a href="/en/llmops/ktransformers-moe-offload-28x-validation/">A $400K rack on 24GB? We reproduced the ktransformers 28x</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="llmops" /><category term="ATSInfer" /><category term="llama.cpp" /><category term="CPU-offloading" /><category term="GPU" /><category term="LLM-serving" /><category term="LLMOps" /><category term="quantization" /><category term="infrastructure" /><summary type="html"><![CDATA[Instead of offloading whole layers or experts, ATSInfer places individual tensors across CPU and GPU. The paper reports running models far beyond VRAM on a single RTX 4090 while pushing decode up to 3.29x. We read arXiv:2607.10183 and unpacked how it works and what it means for serving.]]></summary></entry><entry xml:lang="ko"><title type="html">Kimi Code CLI 뜯어보기: 오픈소스 터미널 에이전트가 ACP로 에디터를 삼키는 법</title><link href="https://thakicloud.github.io/ko/agentops/kimi-code-cli-acp-open-source-agent/" rel="alternate" type="text/html" title="Kimi Code CLI 뜯어보기: 오픈소스 터미널 에이전트가 ACP로 에디터를 삼키는 법" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/ko/agentops/kimi-code-cli-acp-open-source-agent</id><content type="html" xml:base="https://thakicloud.github.io/ko/agentops/kimi-code-cli-acp-open-source-agent/"><![CDATA[<p>지난주 문샷AI가 오픈웨이트 모델 Kimi K3를 공개하면서 코딩 리더보드 1위를 가져갔습니다. 그런데 모델보다 개발자 워크플로에 더 직접 닿는 물건이 조용히 같이 나왔습니다. <strong>Kimi Code CLI</strong>, 문샷이 MIT 라이선스로 공개한 오픈소스 터미널 코딩 에이전트입니다. 링크드인 타임라인에는 “클로드 코드에 없는 기능을 준다”는 소개가 돌았습니다. 저희는 그 문장을 그대로 옮기지 않고 공식 저장소와 문서를 직접 확인했습니다. 결론부터 말하면 소개의 절반은 사실이고 절반은 과장입니다. 그리고 진짜 흥미로운 지점은 홍보 문구가 강조하지 않은 곳에 있었습니다.</p>

<p>이 글은 Kimi Code CLI가 무엇인지, 무엇을 실제로 제공하는지, 그리고 K8s 기반 AI 플랫폼을 운영하는 저희 관점에서 왜 눈여겨볼 만한지를 정리합니다. 특히 Agent Client Protocol이라는 개방 표준이 왜 에이전트 생태계의 판을 바꾸는 조각인지에 지면을 많이 할애했습니다.</p>

<h2 id="kimi-code-cli는-무엇인가">Kimi Code CLI는 무엇인가</h2>

<p>Kimi Code CLI는 터미널에서 도는 에이전트형 코딩 도구입니다. 클로드 코드나 제미나이 CLI, 코덱스 CLI와 같은 계열입니다. 공식 저장소는 <a href="https://github.com/MoonshotAI/kimi-code">MoonshotAI/kimi-code</a>이며, 이전 프로젝트인 <a href="https://github.com/MoonshotAI/kimi-cli">MoonshotAI/kimi-cli</a>가 여기로 진화하면서 기존 세션과 설정이 이어집니다. 두 저장소 모두 문샷 공식이며, 이름이 비슷한 서드파티 프로젝트와 혼동하지 않도록 주의가 필요합니다.</p>

<p>여기서 첫 번째 정정이 필요합니다. 이 CLI의 정식 이름은 “Kimi K3용 CLI”가 아니라 <strong>Kimi Code CLI</strong>입니다. 모델에 종속되지 않는 도구이고, 기본값으로 문샷의 코딩 특화 모델인 Kimi K2.7 Code를 붙여 쓰지만 설정으로 K3를 포함해 다른 모델로도 전환할 수 있습니다. K3는 CLI가 붙일 수 있는 여러 모델 중 하나이지, CLI가 K3 전용으로 만들어진 것은 아닙니다. K3 자체는 문샷이 2026년 7월 16일 공개한 2조8000억 매개변수급 오픈 MoE 모델로, Kimi Delta Attention과 최대 100만 토큰 컨텍스트를 내세웁니다. 이 부분은 CNBC, 블룸버그, 포브스 등 주요 매체가 함께 보도했습니다.</p>

<p>전체 그림을 먼저 세워두면 이후 세부가 훨씬 잘 붙습니다. Kimi Code CLI가 한쪽으로는 MCP 클라이언트로 도구와 데이터에 연결되고, 다른 한쪽으로는 ACP 서버로 에디터에 연결되는 이중 역할이 핵심입니다.</p>

<pre><code class="language-mermaid">flowchart TB
    subgraph EDITOR["개발자 에디터 (ACP 클라이언트)"]
        ZED["Zed"]
        JB["JetBrains 계열"]
        VSC["VS Code / Neovim"]
    end
    ACP["Agent Client Protocol&lt;br/&gt;JSON-RPC over stdio"]
    subgraph CLI["Kimi Code CLI (에이전트 코어)"]
        MAIN["메인 에이전트&lt;br/&gt;대화 히스토리 유지"]
        SUB["서브에이전트&lt;br/&gt;coder · explore · plan&lt;br/&gt;각자 격리 컨텍스트"]
    end
    MODEL["모델 계층&lt;br/&gt;Kimi K2.7 Code / K3&lt;br/&gt;또는 OpenAI 호환 엔드포인트"]
    subgraph MCP["MCP 서버 (도구 · 데이터)"]
        T1["Context7"]
        T2["Chrome DevTools"]
        T3["사내 커넥터"]
    end

    EDITOR --&gt; ACP
    ACP --&gt;|kimi acp| MAIN
    MAIN --&gt; SUB
    MAIN --&gt;|추론 요청| MODEL
    SUB --&gt;|추론 요청| MODEL
    MAIN --&gt;|도구 호출| MCP
</code></pre>

<h2 id="서브에이전트-컨텍스트를-나눠서-메인을-깨끗하게">서브에이전트: 컨텍스트를 나눠서 메인을 깨끗하게</h2>

<p>Kimi Code CLI는 세 종류의 내장 서브에이전트를 제공합니다. <strong>coder</strong>는 파일을 읽고 쓰고 명령을 실행해 실제 변경을 반영하는 범용 엔지니어링 담당입니다. <strong>explore</strong>는 읽기 전용으로 코드베이스를 훑는 탐색 담당입니다. <strong>plan</strong>은 셸 명령 없이 구현 계획과 아키텍처 설계만 내놓는 담당입니다. 이 구분은 공식 문서 <a href="https://moonshotai.github.io/kimi-code/en/customization/agents.html">Agents and Sub-Agents</a>에 명시되어 있습니다.</p>

<p>핵심은 이름이 아니라 컨텍스트 격리입니다. 각 서브에이전트는 완전히 독립된 컨텍스트 윈도우를 가지며, 메인 에이전트가 명시적으로 넘긴 작업 설명만 봅니다. 메인의 대화 히스토리는 서브에이전트에 노출되지 않고, 서브에이전트가 돌리는 중간 추론과 도구 호출 로그도 메인 히스토리에 섞이지 않습니다. 서브에이전트는 최종 결론만 반환합니다. 긴 세션에서 메인 컨텍스트가 로그로 부풀지 않고 얇게 유지되는 이유가 여기 있습니다. 백그라운드 실행과 병렬 실행도 지원해서, 여러 탐색 작업을 동시에 돌리고 완료되면 자동으로 결과가 돌아옵니다.</p>

<p>이 패턴은 저희에게 낯설지 않습니다. 이 블로그를 운영하는 내부 오케스트레이션 하네스도 탐색은 저비용 서브에이전트에 위임하고 요약만 회수해 메인 컨텍스트를 보호합니다. 컨텍스트 위생이 곧 비용이자 품질이라는 원칙은 도구가 달라도 같습니다.</p>

<h2 id="mcp-json을-손으로-고치지-않는-설정-경험">MCP: JSON을 손으로 고치지 않는 설정 경험</h2>

<p>Model Context Protocol 연동은 두 경로로 관리합니다. 첫째, CLI 서브커맨드입니다. <code class="language-plaintext highlighter-rouge">kimi mcp add</code>, <code class="language-plaintext highlighter-rouge">kimi mcp list</code>, <code class="language-plaintext highlighter-rouge">kimi mcp remove</code>, <code class="language-plaintext highlighter-rouge">kimi mcp authorize</code>로 서버를 다룹니다. 예를 들어 HTTP 트랜스포트로 문서 검색 서버를 붙이거나, stdio 트랜스포트로 브라우저 자동화 서버를 붙일 수 있습니다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># HTTP 트랜스포트 (OAuth 옵션 지원)</span>
kimi mcp add <span class="nt">--transport</span> http context7 https://mcp.context7.com/mcp

<span class="c"># stdio 트랜스포트로 로컬 프로세스 연결</span>
kimi mcp add <span class="nt">--transport</span> stdio chrome-devtools <span class="nt">--</span> npx chrome-devtools-mcp@latest
</code></pre></div></div>

<p>둘째, TUI 안에서 쓰는 대화형 슬래시 명령 <code class="language-plaintext highlighter-rouge">/mcp-config</code>입니다. JSON 설정 파일을 직접 편집하지 않고 서버를 추가, 수정, 인증할 수 있습니다. <code class="language-plaintext highlighter-rouge">/mcp</code>는 현재 연결된 서버와 로드된 도구 목록을 보여줍니다. 링크드인 소개가 강조한 “JSON을 직접 수정할 필요가 없다”는 부분은 사실입니다. 다만 이 편의성 자체가 클로드 코드에 없는 것은 아닙니다. 이 지점은 뒤에서 다시 정리합니다. 관련 문서는 <a href="https://moonshotai.github.io/kimi-cli/en/customization/mcp.html">MCP 설정</a>에 있습니다.</p>

<h2 id="agent-client-protocol-이-도구에서-가장-중요한-조각">Agent Client Protocol: 이 도구에서 가장 중요한 조각</h2>

<p>여기가 이 글에서 가장 흥미로운 부분입니다. Agent Client Protocol, 줄여서 ACP는 Zed 에디터 팀이 만든 개방 표준입니다. Apache 라이선스이고, JSON-RPC 2.0을 stdio 위에서 주고받습니다. 에디터가 에이전트를 자식 프로세스로 띄우고 표준 입출력으로 통신하는 방식으로, 전송 메커니즘 자체는 언어 서버 프로토콜과 동일합니다.</p>

<p>비유가 이해를 크게 돕습니다. LSP가 등장하기 전에는 에디터마다 언어마다 별도 통합을 만들어야 했습니다. LSP는 이 M 곱하기 N 문제를 M 더하기 N으로 바꿨습니다. 에디터 하나가 표준만 구현하면 누가 만든 언어 서버든 그 덕을 봅니다. ACP는 정확히 같은 일을 에이전트에 합니다. 에디터 하나가 ACP를 구현하면 누가 만든 에이전트든 그 에디터에 표준 방식으로 꽂힙니다. 이 개념은 <a href="https://zed.dev/acp">Zed의 ACP 소개</a>와 <a href="https://blog.marcnuri.com/agent-client-protocol-acp-introduction">마크 누리의 해설 글</a>에서 확인할 수 있습니다.</p>

<p>MCP와 헷갈리기 쉬운데 방향이 반대입니다. MCP는 에이전트에서 도구와 데이터로 향합니다. 이때 에이전트가 MCP 클라이언트입니다. ACP는 에디터에서 에이전트로 향합니다. 이때 에이전트가 ACP 서버이고 에디터가 ACP 클라이언트입니다. 같은 에이전트가 한쪽으로는 MCP 클라이언트, 다른 쪽으로는 ACP 서버 역할을 동시에 맡습니다. 앞의 다이어그램이 이 이중 역할을 그린 이유입니다.</p>

<p>Kimi Code CLI는 <code class="language-plaintext highlighter-rouge">kimi acp</code> 서브커맨드로 이 프로토콜을 별도 설치 없이 네이티브로 지원합니다. Zed는 네이티브로, JetBrains는 플러그인으로 연결되고, Zed의 ACP 레지스트리 기준으로는 여러 에디터 통합이 이미 올라와 있습니다. 개발자는 익숙한 에디터를 떠나지 않고 그 안에서 Kimi 세션을 몰 수 있습니다.</p>

<h2 id="이미지와-비디오-입력-정확히-어디까지-되나">이미지와 비디오 입력, 정확히 어디까지 되나</h2>

<p>링크드인 소개는 “화면 캡처를 그대로 입력으로 전달할 수 있다”고 적었습니다. 여기에 정정이 필요합니다. 문샷이 정면에 내세우는 기능은 정적 스크린샷이 아니라 <strong>화면 녹화 영상 입력</strong>입니다. 저장소 설명은 화면 녹화나 데모 클립을 채팅에 떨어뜨리면 에이전트가 말로 설명하기 어려운 동작을 직접 보고 이해한다고 표현합니다. 물론 CLI 입력창에서 이미지 붙여넣기도 지원하며, 기본 모델 Kimi K2.7 Code가 4억 매개변수 비전 인코더 MoonViT를 갖춘 네이티브 멀티모달이라 텍스트, 이미지, 비디오를 모두 받습니다. 다만 커스텀 모델을 붙일 때는 해당 모델의 modalities에 이미지 지원을 명시해야 정상 동작합니다. 요약하면 이미지 입력이 되긴 하지만, 진짜 차별 포인트로 홍보되는 것은 영상 입력이며 “스크린샷”이라는 표현은 다소 부정확합니다.</p>

<h2 id="설치는-실제로-세-단계">설치는 실제로 세 단계</h2>

<p>설치 흐름은 소개대로 간결합니다. 아래 명령은 <a href="https://moonshotai.github.io/kimi-cli/en/guides/getting-started.html">공식 시작 가이드</a> 기준이며, 저희 사내 샌드박스가 해당 배포 도메인에 접근 권한이 없어 직접 실행 로그는 남기지 않았습니다. 따라서 벤치마크 수치는 만들지 않고, 검증된 명령만 옮깁니다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># 1) 설치 스크립트 실행 (uv를 함께 설치)</span>
curl <span class="nt">-LsSf</span> https://code.kimi.com/install.sh | bash

<span class="c"># 2) 프로젝트 디렉터리에서 실행</span>
kimi

<span class="c"># 3) 인증 설정</span>
/login
</code></pre></div></div>

<p>macOS는 <code class="language-plaintext highlighter-rouge">brew install kimi-code</code>, 윈도우는 파워셸 스크립트도 제공합니다. 소스에서 개발하려면 Node 24.15 이상과 pnpm이 필요합니다. 라이선스는 MIT라서, 코드를 열어 보고 포크하고 사내 배포하는 데 제약이 적습니다.</p>

<h2 id="모델과-프로바이더-개방성">모델과 프로바이더 개방성</h2>

<p>컨텍스트 길이는 K2.6 계열이 최대 25만6000 토큰이고, K3는 문샷 마케팅 기준 최대 100만 토큰입니다. 더 중요한 것은 프로바이더 개방성입니다. <code class="language-plaintext highlighter-rouge">~/.kimi-code/config.toml</code>에서 OpenAI 호환 엔드포인트, Anthropic API 키, 구글 GenAI나 Vertex AI까지 다중 프로바이더로 등록할 수 있습니다. CLI가 특정 모델에 락인되지 않는다는 뜻입니다. 서드파티 추론 모델의 reasoning_content 필드도 자동 처리합니다. 관련 문서는 <a href="https://moonshotai.github.io/kimi-cli/en/configuration/providers.html">Providers and models</a>에 있습니다.</p>

<h2 id="클로드-코드에는-없는-기능인가-정직한-비교">클로드 코드에는 없는 기능인가: 정직한 비교</h2>

<p>가장 널리 퍼진 소개 문구는 “클로드 코드에는 없는 기능을 제공한다”였습니다. 검증 결과 이 프레이밍은 대부분 과장입니다.</p>

<p>서브에이전트와 격리 컨텍스트는 클로드 코드도 서브에이전트 기능으로 같은 방식을 제공합니다. MCP도 클로드 코드가 stdio, SSE, HTTP 전송을 이미 성숙하게 지원합니다. 이미지 붙여넣기도 클로드 코드에 있습니다. 이 세 가지는 차별점이 아닙니다.</p>

<p>진짜 차이는 두 곳에 있습니다. 첫째, ACP 지원 방식입니다. Kimi Code CLI는 <code class="language-plaintext highlighter-rouge">kimi acp</code> 서브커맨드로 CLI 자체에 ACP를 1차 기능으로 내장합니다. 반면 클로드 코드는 Zed가 만든 별도 어댑터 패키지를 거쳐 베타로 연결됩니다. 사용자 입장에서 전자는 도구를 켜면 바로 되고, 후자는 브릿지를 하나 더 얹어야 합니다. 둘째, 모델 개방성입니다. Kimi는 오픈웨이트 K 시리즈에 멀티 프로바이더 전환까지 열려 있는 반면, 클로드 코드는 앤트로픽 모델 전용입니다. 여기서 셀프호스팅 가능성이라는 세 번째 차이가 파생됩니다. Kimi는 오픈소스 CLI에 오픈웨이트 모델이라 온프렘 서빙이 가능하지만, 클로드 코드는 CLI는 열려 있어도 모델이 API 전용입니다. 관련 근거는 <a href="https://zed.dev/blog/claude-code-via-acp">Zed의 클로드 코드 ACP 베타 글</a>에서 확인할 수 있습니다.</p>

<h2 id="thakicloud-제품-적용-시사점">ThakiCloud 제품 적용 시사점</h2>

<p>이 주제는 에이전트 도구인 동시에 개방 모델과 온프렘이라는 인프라 축을 건드립니다. 그래서 두 렌즈를 함께 씁니다.</p>

<p>Paxis 렌즈로 보면 Kimi Code CLI의 구조가 저희 제품의 설계 방향과 상당히 겹칩니다. Paxis는 ThakiCloud의 Agent-Native Cloud 제어 평면으로, 스킬과 도구와 정책과 감사 로그를 일급 리소스로 다룹니다. Kimi의 coder, explore, plan 서브에이전트가 격리 컨텍스트에서 병렬로 도는 방식은 Paxis의 스킬 하네스가 960개 넘는 스킬을 BM25로 선택해 격리 샌드박스에서 실행하는 방식과 같은 철학을 공유합니다. 특히 ACP는 벤더 중립 표준이라는 점에서 Paxis에 직접적인 기회입니다. 저희가 배포하는 어떤 에이전트든, 자체 파인튜닝 모델을 얹은 것까지 포함해서, ACP를 구현하면 고객의 Zed나 JetBrains 같은 개발 에디터에 표준 방식으로 꽂힐 수 있습니다. MCP 커넥터로 데이터에 연결하고 ACP로 에디터에 연결하는 이중 표준 조합은 정확히 저희가 지향하는 통합 그림입니다.</p>

<p>ai-platform 렌즈로 보면 개방성이 곧 배포 자유입니다. 오픈웨이트 K 시리즈를 저희 클러스터의 Kueue GPU 스케줄링과 vLLM 서빙 위에 올리고, CLI를 사내 엔드포인트로 라우팅하면 외부 API 의존이나 데이터 반출 없이 사내 코딩 에이전트를 구축할 수 있습니다. 코드가 외부로 나가면 안 되는 금융이나 공공 영역, 그리고 국정원 요구사항 같은 온프렘 보안 요건과 정합합니다. 능력이 흔해지고 싸질수록 기업이 실제로 지불하는 것은 통제된 실행 환경이라는 관점은 저희가 이전 글에서도 다룬 주제입니다. Kimi Code CLI는 그 실행 계층을 오픈소스로 열어 놓았다는 점에서 의미가 있습니다.</p>

<h2 id="한계-및-반론">한계 및 반론</h2>

<p>몇 가지는 냉정하게 봐야 합니다. 첫째, 서드파티 딥다이브에서 언급되는 내부 엔진 명칭이나 계층 구조는 공식 문서에 없는 표현이라 리버스 엔지니어링일 가능성이 있습니다. 사실로 인용하기 전에 공식 문서를 기준으로 삼는 편이 안전합니다. 둘째, ACP 경로가 다른 연결 방식보다 응답 품질이 낫다는 커뮤니티 보고가 있지만 벤치마크가 아니라 체감입니다. 검증된 수치가 아닙니다. 셋째, 오픈웨이트라 해도 2조8000억 매개변수급 모델을 온프렘에서 실제로 서빙하려면 상당한 GPU 자원이 필요합니다. 개방성이 곧 손쉬운 셀프호스팅을 뜻하지는 않습니다. 작은 팀에는 API 경로가 여전히 현실적입니다. 넷째, 도구 생태계의 성숙도와 안정성은 클로드 코드나 코덱스 CLI가 앞서 있을 수 있습니다. 오픈소스라는 사실이 곧 프로덕션 준비 완료를 의미하지는 않습니다.</p>

<p>그럼에도 개방 표준 위에서 에이전트와 에디터가 느슨하게 결합되는 방향은 분명한 흐름입니다. 특정 벤더의 CLI에 종속되지 않고, 모델과 에디터를 각각 갈아 끼울 수 있는 세계가 개발자에게 더 유리합니다. Kimi Code CLI는 그 세계를 앞당기는 조각 중 하나입니다.</p>

<h2 id="출처">출처</h2>

<ul>
  <li><a href="https://github.com/MoonshotAI/kimi-code">MoonshotAI/kimi-code (공식 저장소)</a></li>
  <li><a href="https://github.com/MoonshotAI/kimi-cli">MoonshotAI/kimi-cli (이전 저장소)</a></li>
  <li><a href="https://moonshotai.github.io/kimi-cli/en/guides/getting-started.html">Kimi Code CLI 시작 가이드</a></li>
  <li><a href="https://moonshotai.github.io/kimi-code/en/customization/agents.html">Agents and Sub-Agents 문서</a></li>
  <li><a href="https://moonshotai.github.io/kimi-cli/en/customization/mcp.html">MCP 설정 문서</a></li>
  <li><a href="https://zed.dev/acp">Zed - Agent Client Protocol</a></li>
  <li><a href="https://blog.marcnuri.com/agent-client-protocol-acp-introduction">ACP: The LSP for AI Coding Agents</a></li>
  <li><a href="https://zed.dev/blog/claude-code-via-acp">Zed - Claude Code via ACP (베타)</a></li>
  <li><a href="https://www.marktechpost.com/2026/07/16/moonshot-ai-releases-kimi-k3-a-2-8-trillion-parameter-open-moe-model-with-kimi-delta-attention-and-1m-context/">MarkTechPost - Kimi K3 공개</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="agentops" /><category term="agentops" /><category term="kimi" /><category term="moonshot" /><category term="coding-agent" /><category term="mcp" /><category term="agent-client-protocol" /><category term="paxis" /><category term="thakicloud" /><summary type="html"><![CDATA[문샷AI가 Kimi K3와 함께 공개한 오픈소스 코딩 CLI를 실제 문서와 저장소 기준으로 분석합니다. coder/explore/plan 서브에이전트, 대화형 MCP 설정, 그리고 진짜 차별점인 Agent Client Protocol 네이티브 지원까지, '클로드 코드에 없는 기능'이라는 홍보 문구가 어디까지 사실인지 검증합니다.]]></summary></entry><entry xml:lang="ko"><title type="html">250억 원짜리 청구서가 남긴 질문: AI 에이전트의 비용은 왜 보이지 않는가</title><link href="https://thakicloud.github.io/ko/llmops/agent-cost-observability-billing-crisis/" rel="alternate" type="text/html" title="250억 원짜리 청구서가 남긴 질문: AI 에이전트의 비용은 왜 보이지 않는가" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/ko/llmops/agent-cost-observability-billing-crisis</id><content type="html" xml:base="https://thakicloud.github.io/ko/llmops/agent-cost-observability-billing-crisis/"><![CDATA[<p>이 글은 조직에 Claude Code나 AI 에이전트를 도입하려는 플랫폼·인프라 담당자, 그리고 다음 달 AI 청구서를 설명해야 하는 재무·구매 담당자를 위해 썼습니다. 결론을 먼저 말씀드리면, 최근 한 달 동안 쏟아진 AI 비용 뉴스는 “AI가 비싸다”는 이야기가 아닙니다. 진짜 문제는 <strong>청구서가 무엇에 대한 것인지 설명해 주지 않는다</strong>는 데 있습니다. 한 번의 사용자 요청이 수십에서 수백 번의 모델 호출과 도구 실행, 그리고 실패 시 자동 재시도로 번지는 에이전트 구조에서, 최종 금액만으로는 어느 루프에서 돈이 샜는지 알 수가 없습니다. 저희는 이 관측 공백이야말로 지금 시장이 겪는 통증의 핵심이라고 봅니다.</p>

<h2 id="개요">개요</h2>

<p>2026년 6월 말부터 7월까지의 뉴스를 한 줄로 요약하면 이렇습니다. 프런티어 모델을 모든 작업에 무차별로 쓰면 비용을 감당하기 어렵고, 그 비용이 어디서 발생했는지조차 추적하기 어렵다는 것입니다. 사건은 극적인 형태로 드러났습니다. 국내 한 사용자에게 처음 약 25억 원, 이후 약 250억 원의 결제가 시도됐습니다. 다행히 카드 한도 초과로 실제 출금은 없었지만, 비정상적인 금액이 표시 오류를 넘어 카드 승인 요청 단계까지 반복해서 전달됐다는 점이 사건의 무게를 다르게 만듭니다.</p>

<p>같은 시기, 다른 층위의 뉴스도 이어졌습니다. AI 비용 감사 업체가 기업 60곳의 청구서를 검토해 상당한 과다 청구를 주장했고, 여러 대기업이 프런티어 모델 사용을 조건에 따라 저가 모델로 분산하기 시작했으며, 미국·유럽 기업들이 비용을 이유로 중국 오픈웨이트 모델로 옮겨 갔다는 보도가 나왔습니다. 흥미롭게도 모델 공급사 스스로도 “모든 작업에 최고 성능 모델을 오래 돌리는 방식은 지속 가능하지 않다”는 취지로 대응하기 시작했습니다. 여러 방향의 뉴스가 같은 지점을 가리키고 있었습니다.</p>

<h2 id="지난-한-달-무슨-일이-있었나">지난 한 달, 무슨 일이 있었나</h2>

<p>가장 먼저 눈에 띈 것은 청구 시스템의 신뢰성 문제였습니다. 지디넷코리아 보도에 따르면, 국내 대학생 이용자에게 발생한 청구액은 약 166만 달러에서 시작해 열 배 규모인 약 1,662만 달러까지 늘어났습니다. 앤트로픽은 이후 답변에서 자동 충전 금액이 비정상적으로 높게 설정된 오류였다고 설명했지만, 당사자는 자동 충전 기능을 설정한 적이 없다고 밝혀 설정이 왜 생성됐는지는 여전히 명확하지 않습니다. 사용자가 여러 부서에 열다섯 통 넘는 메일을 보낸 뒤 나흘이 지나서야 자동 응답을 받았다는 대목은 기술 오류보다 대응 체계의 공백을 더 선명하게 보여 줍니다.</p>

<p>두 번째는 에이전트 비용의 관측 가능성 문제였습니다. AI타임스와 디인포메이션 보도에 따르면, 비용 감사 스타트업 보디트(Vaudit)는 기업 60곳의 약 3,400만 달러 청구서를 검토해 약 170만 달러를 과다 청구로 판단했다고 밝혔습니다. 검토 대상의 상당 부분은 Claude Code 사용 내역이었고, 파나소닉·HP·혼다 등이 고객사로 언급됐습니다. 이 업체가 주장한 유형은 저가 모델을 썼는데 고가 모델 요금으로 기록되거나, 작업을 완료하지 못했는데 비용이 발생하거나, 오류 후 자동 재시도가 반복되는 이른바 <strong>리트라이 스톰(retry storm)</strong>이었습니다. 여기서 짚어 둘 점이 두 가지 있습니다. 첫째, 앤트로픽은 완료되지 않은 요청이나 오류 응답에 비용을 청구하지 않으며 광범위한 과다 청구 증거도 없다고 반박했습니다. 둘째, 보디트는 환불 성공액의 일부를 수수료로 받는 상업적 감사 업체이므로, 이 수치는 독립 회계감사가 아니라 한쪽 당사자의 조사 결과로 읽는 것이 정확합니다. 즉 지금은 감사 업체의 주장과 공급사의 부인이 맞서는 국면입니다.</p>

<p>세 번째는 시장의 반응이었습니다. 디인포메이션은 기업들이 단순 분류·요약·변환 같은 작업은 저가 모델로, 복잡한 코딩·에이전트 작업은 프런티어 모델로, 반복적인 대량 작업은 오픈웨이트 또는 자체 호스팅 모델로 분리하기 시작했다고 보도했습니다. 파이낸셜타임스는 도어대시·지멘스·에어비앤비 등이 비용 절감을 위해 딥시크나 문샷 계열 모델을 도입했다고 전했습니다. 비즈니스인사이더 보도에서는 앤트로픽의 플랫폼 책임자들조차 부서별로 제각각 도입하는 이른바 <strong>섀도 IT</strong> 때문에 일부 기업의 AI 비용이 폭증했다고 인정하면서도, 사용 중단이나 일괄 예산 상한이 아니라 작업별 모델 선택과 조직 차원의 중앙 비용 관리가 필요하다고 주장했습니다. 요금 정책 자체도 자주 바뀌었습니다. 최신 고성능 모델의 구독 포함 여부와 종량제 전환 시점이 여러 차례 조정됐고, 프로모션 종료 시점이 반복해서 연장됐습니다. 실제로 Claude Fable 5 무료 제공이 7월 19일까지 연장됐다는 보도도 있었습니다. 성능보다 다음 달 비용을 예측하기 어렵다는 점이 구매 담당자에게는 더 큰 골칫거리였습니다.</p>

<h2 id="에이전트-비용은-왜-관측되지-않는가">에이전트 비용은 왜 관측되지 않는가</h2>

<p>세 갈래의 뉴스를 관통하는 공통 원인은 결국 하나입니다. 에이전트 워크로드에서는 사용자가 보는 것과 청구서가 기록하는 것 사이의 거리가 너무 멀어졌습니다. 전통적인 API 호출은 요청 한 번에 응답 한 번, 비용 한 줄이었습니다. 반면 코딩 에이전트나 에이전트 SDK는 한 번의 지시가 계획 수립, 도구 호출, 파일 편집, 검증, 실패 시 재시도로 확장됩니다. 이 확장은 사용자에게 보이지 않는 곳에서 일어나고, 청구서에는 그 총합만 한 줄로 찍힙니다.</p>

<pre><code class="language-mermaid">flowchart TB
    U["사용자 요청 1건"] --&gt; P["에이전트 계획 수립"]
    P --&gt; L["실행 루프"]
    L --&gt; T["도구 호출 · 모델 호출&lt;br/&gt;수십~수백 회"]
    T --&gt; R{"성공 여부"}
    R --&gt;|"실패"| RS["자동 재시도&lt;br/&gt;(리트라이 스톰)"]
    RS --&gt; T
    R --&gt;|"성공"| ACC["토큰 · 캐시 · tool call&lt;br/&gt;누적 집계"]
    ACC --&gt; INV["청구서: 최종 금액 한 줄"]
    INV -.관측 공백.-&gt; U
</code></pre>

<p>이 구조에서 비용이 새는 지점은 대부분 사용자의 시야 밖에 있습니다. 재시도 루프가 조용히 돌면서 호출 수를 부풀리고, 중간 클라우드 사업자를 거치면서 실제 모델 사용량과 최종 청구 내역이 어긋나며, 자동 충전 같은 설정 하나가 잘못되면 카드 승인 단계까지 비정상 금액이 흘러갑니다. 세 뉴스는 서로 다른 사건처럼 보이지만, 전부 이 관측 공백의 다른 얼굴입니다. 그래서 개인별 월 한도를 거는 것만으로는 부족합니다. 필요한 것은 모델별 비용, 세션별 토큰, 캐시 토큰, 도구 호출 횟수, 실패와 재시도 비용, 일별 이상 증가율을 <strong>호출이 일어나는 그 순간에</strong> 중앙에서 붙잡는 계측 계층입니다. 관측이 없으면 통제도 없고, 통제가 없으면 청구서는 늘 사후에 놀라는 서류가 됩니다.</p>

<h2 id="thakicloud-제품-적용-시사점">ThakiCloud 제품 적용 시사점</h2>

<p>이 문제는 ThakiCloud가 운용하는 두 제품이 각기 다른 각도에서 겨냥하는 지점입니다. 인프라 관점과 에이전트 관점이 서로를 보완하기 때문에, 이번 주제에는 두 렌즈를 함께 씁니다.</p>

<p><strong>ai-platform 렌즈, 반복 워크로드는 소유가 답입니다.</strong> 시장이 도달한 결론은 명료합니다. 쉬운 작업까지 프런티어 모델로 처리하면 비용이 감당되지 않고, 반복적인 대량 작업은 오픈웨이트 모델을 자체 호스팅하는 편이 경제적입니다. ThakiCloud의 ai-platform은 바로 이 지점을 위한 K8s 기반 AI/ML 인프라입니다. Kueue로 GPU를 큐잉해 활용률을 끌어올리고, vLLM으로 오픈웨이트 모델을 서빙하며, 멀티테넌트 격리로 부서별 사용량을 분리해 과금합니다. 종량제 API가 예측 불가능한 청구서를 만든다면, 자체 호스팅은 고정된 GPU 비용 위에서 사용량이 늘어도 단가가 튀지 않는 구조를 만듭니다. 요금 정책이 수시로 바뀌는 외부 종량제와 달리, 온프렘·소버린 배치는 비용의 예측 가능성 자체를 자산으로 돌려줍니다. 데이터가 외부로 나가지 않는다는 점은 국내 규제·보안 요구가 큰 조직에는 별도의 가치가 됩니다.</p>

<p><strong>Paxis 렌즈, 에이전트의 모든 행동을 감사 가능하게 만듭니다.</strong> 관측 공백의 핵심은 에이전트 루프였고, 이것은 정확히 Paxis가 다루는 영역입니다. Paxis는 ai-platform 위에서 도는 ThakiCloud의 Agent-Native Cloud 제어 평면으로, Skills·Tools·Policies·Audit Logs를 일급 리소스로 취급합니다. 에이전트가 어떤 스킬을 어떤 도구로 몇 번 호출했는지, 어느 샌드박스에서 실행했는지가 전부 감사 로그로 남습니다. 이 구조에서는 리트라이 스톰이 조용히 청구서를 부풀리는 대신, 재시도 루프가 감사 로그에 그대로 드러나고 정책 게이트가 임계를 넘는 호출을 차단합니다. 960개가 넘는 스킬을 BM25로 선택해 격리 샌드박스에서 실행하고 모든 행동을 정책과 감사로 통과시키는 설계는, “청구서만 보고 어느 루프에서 비용이 났는지 확인하기 어렵다”는 바로 그 문제에 대한 구조적 답입니다. 저비용 서빙(ai-platform)이 에이전트를 경제적으로 만들고, 행동 단위 관측(Paxis)이 그 경제성을 예측 가능하게 만듭니다. 두 렌즈는 이렇게 맞물립니다.</p>

<h2 id="한계-및-반론">한계 및 반론</h2>

<p>균형을 위해 반대편도 분명히 해 두겠습니다. 첫째, 프런티어 모델이 무조건 낭비인 것은 아닙니다. 월스트리트저널 보도에 따르면 쇼피파이 같은 기업은 복잡한 코딩과 다단계 에이전트 작업에서 프런티어 모델이 엔지니어의 시간을 아껴 준다면 높은 가격도 정당화될 수 있다고 봅니다. 반대로 스포티파이나 트윌리오는 소폭의 성능 향상이 추가 비용을 정당화하는지 신중하게 따지고 있습니다. 즉 답은 “프런티어를 버리라”가 아니라 “작업 난이도에 따라 나누라”입니다. 자체 호스팅도 만능이 아닙니다. 최고 난도의 추론이 필요한 작업까지 오픈웨이트로 내리면 품질이 떨어지고, GPU 운영·모델 업데이트·보안 패치라는 새로운 운영 부담이 생깁니다.</p>

<p>둘째, 이 글에 인용한 과다 청구 수치는 확정된 사실이 아닙니다. 보디트의 주장은 상업적 감사 업체의 발표이고 앤트로픽은 이를 부인했으므로, 현재로서는 양측의 입장이 맞서는 상태로 읽는 것이 정확합니다. 250억 원 청구 사건 역시 실제 출금은 이뤄지지 않았고, 자동 충전 설정이 왜 생성됐는지에 대한 기술적 설명은 아직 공개되지 않았습니다. 저희가 이 뉴스에서 끌어내는 결론은 특정 공급사를 겨냥하는 것이 아니라, 에이전트 시대에는 어느 공급사를 쓰든 비용 관측성과 거버넌스를 사용자 쪽에서 확보해야 한다는 원칙입니다. 좋은 모델을 고르는 문제와 그 모델을 다스리는 문제는 별개이고, 최근 한 달의 뉴스는 후자가 그동안 비어 있었다는 사실을 드러냈을 뿐입니다.</p>

<h2 id="출처">출처</h2>

<ul>
  <li><a href="https://zdnet.co.kr/view/?no=20260709165452">지디넷코리아, “250억원 결제 요청 받은 국내 이용자…앤트로픽 빌링 오류 논란” (2026-07-09)</a></li>
  <li><a href="https://zdnet.co.kr/view/?no=20260716093004">지디넷코리아, “250억원 청구한 앤트로픽, 알고 보니 자동 충전 설정 오류” (2026-07-16)</a></li>
  <li><a href="https://www.aitimes.com/news/articleView.html?idxno=212155">AI타임스, “앤트로픽, ‘AI 비용 과다 청구’ 논란…실패한 작업도 돈 받았다” (2026-06)</a></li>
  <li><a href="https://www.theinformation.com/titv/fedld">The Information, 기업의 AI 비용 통제와 모델 분산 도입 보도 (2026-06-23)</a></li>
  <li><a href="https://www.ft.com/content/9c8ff45b-7c20-4c2e-93c9-c52339ffdcee">Financial Times, “Companies turn to Chinese AI models to cut costs” (2026-07)</a></li>
  <li><a href="https://www.businessinsider.com/anthropic-ai-costs-responses-routers-2026-7">Business Insider, “Anthropic Official Warns Against ‘Wrong’ AI Cost Response” (2026-07-15)</a></li>
  <li><a href="https://www.wsj.com/cio-journal/meet-the-companies-shelling-out-for-top-ai-models-e1fe3375">The Wall Street Journal, “Meet the Companies Shelling Out for Top AI Models” (2026-07)</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="llmops" /><category term="LLMOps" /><category term="FinOps" /><category term="에이전트비용" /><category term="비용관측성" /><category term="모델라우팅" /><category term="self-hosting" /><category term="Paxis" /><category term="AI인프라" /><summary type="html"><![CDATA[국내 사용자에게 250억 원이 청구된 사건부터 기업 60곳의 과다 청구 의혹까지, 최근 한 달의 AI 비용 뉴스는 하나의 공백을 가리킵니다. 한 번의 요청이 수백 번의 모델 호출로 번지는 에이전트 시대에, 청구서는 더 이상 무엇에 돈을 냈는지 설명해 주지 않습니다. 이 관측 공백을 어떻게 메울지 정리했습니다.]]></summary></entry><entry xml:lang="ko"><title type="html">24GB 그래픽카드로 122B 모델을? llama.cpp를 텐서 단위로 쪼갠 ATSInfer를 뜯어봤습니다</title><link href="https://thakicloud.github.io/ko/llmops/atsinfer-hybrid-cpu-gpu-tensor-scheduling/" rel="alternate" type="text/html" title="24GB 그래픽카드로 122B 모델을? llama.cpp를 텐서 단위로 쪼갠 ATSInfer를 뜯어봤습니다" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/ko/llmops/atsinfer-hybrid-cpu-gpu-tensor-scheduling</id><content type="html" xml:base="https://thakicloud.github.io/ko/llmops/atsinfer-hybrid-cpu-gpu-tensor-scheduling/"><![CDATA[<p>이 글은 소비자용 GPU 한 장으로 거대 모델을 자체 서빙할지 저울질하는 엔지니어, 그리고 “24GB로 120B를 돌린다”는 요즘 트윗을 어디까지 믿어야 할지 판단해야 하는 인프라 담당자를 위해 썼습니다. 결론부터 말하면, 난징대학교 연구진이 공개한 ATSInfer(arXiv:2607.10183)의 핵심 아이디어는 단순하면서도 설득력이 있습니다. 지금까지의 오프로딩이 “레이어” 또는 “전문가(expert)” 단위로 뭉텅이째 CPU와 GPU를 오갔다면, ATSInfer는 그 단위를 <strong>텐서 하나하나</strong>까지 쪼갭니다. 다만 화제가 된 “최대 3.29배”라는 숫자는 몇 가지 전제 위에 서 있고, 아직 코드가 공개되지 않았다는 점도 함께 짚겠습니다. 저희는 RTX 4090과 120B급 모델을 이 자리에서 재현하지는 못했으므로, 이 글의 모든 수치는 <strong>논문이 보고한 값</strong>임을 먼저 분명히 해 둡니다.</p>

<h2 id="개요">개요</h2>

<p>로컬 LLM을 돌려 본 사람이라면 한 번쯤 마주치는 벽이 있습니다. 모델 가중치가 GPU 메모리보다 크면, 남는 부분은 CPU 메모리로 내려야 합니다. llama.cpp의 <code class="language-plaintext highlighter-rouge">-ngl</code>(GPU에 올릴 레이어 수) 플래그가 바로 그 일을 합니다. 문제는 이 방식이 <strong>레이어 단위</strong>로만 자른다는 점입니다. 한 레이어 안에는 어텐션 가중치, FFN 가중치, 정규화 파라미터처럼 성격이 전혀 다른 텐서들이 섞여 있는데, 이들을 통째로 “GPU에 올리거나 / CPU에 두거나” 둘 중 하나로만 처리합니다.</p>

<p>이 뭉텅이 배치가 왜 손해인지는 간단합니다. 같은 1GB를 VRAM에 올려도, 어떤 텐서는 GPU에서 10배 빨라지고 어떤 텐서는 2배밖에 안 빨라집니다. VRAM은 희소 자원인데, 뭉텅이로 자르면 “GB당 이득이 큰 텐서”를 골라 담을 수가 없습니다. ATSInfer는 이 지점을 정확히 겨냥합니다. 텐서마다 CPU와 GPU에서의 성능을 프로파일링해, <strong>VRAM 1GB당 속도 이득이 가장 큰 텐서부터</strong> 채워 넣습니다. 최근 화제였던 ktransformers가 MoE 모델의 전문가를 CPU로 내리는 “전문가 단위” 트릭이었다면(<a href="/ko/llmops/ktransformers-moe-offload-28x-validation/">관련 글: ktransformers의 28배를 재현해봤습니다</a>), ATSInfer는 그보다 한 단계 더 잘게 쪼갠 “텐서 단위” 일반화라고 볼 수 있습니다. MoE뿐 아니라 밀집(dense) 모델에도 적용된다는 점이 특히 다릅니다.</p>

<h2 id="이-기술은-무엇인가">이 기술은 무엇인가</h2>

<p>ATSInfer는 llama.cpp를 약 1만 5천 줄의 C++로 확장한 하이브리드 CPU-GPU 추론 시스템입니다. 이름 그대로 “자동 텐서 스케줄링(Automated Tensor Scheduling)”이 핵심이며, 세 가지 메커니즘이 맞물려 돌아갑니다.</p>

<pre><code class="language-mermaid">flowchart TB
    A["모델 가중치&lt;br/&gt;(RAM, VRAM 용량 초과)"] --&gt; B["텐서별 성능 프로파일링&lt;br/&gt;GB당 속도 이득 측정"]
    B --&gt; C{"정적 배치 결정&lt;br/&gt;이득 큰 텐서부터 VRAM에"}
    C --&gt;|"고이득 텐서"| D["VRAM 상주"]
    C --&gt;|"저이득 텐서"| E["RAM 상주"]
    D --&gt; F["로드 인식 동적 전송&lt;br/&gt;런타임 부하 따라 승격·강등"]
    E --&gt; F
    F --&gt; G["비동기 CPU-GPU 조율&lt;br/&gt;연산·PCIe 전송 오버랩"]
    G --&gt; H["토큰 출력&lt;br/&gt;prefill · decode"]
</code></pre>

<p><strong>첫째, 정적 텐서 배치(static placement)입니다.</strong> 모델을 올리기 전에, 벤치마크로 각 텐서가 GPU에서 얼마나 빨라지는지를 미리 측정합니다. 그리고 “VRAM 1GB를 썼을 때 가장 많은 속도를 돌려주는” 텐서 순으로 GPU에 배치합니다. 이는 배낭 문제(knapsack)에 가까운 최적화로, 뭉텅이 배치가 놓치던 텐서 간 이질성을 정면으로 활용합니다.</p>

<p><strong>둘째, 로드 인식 동적 전송(load-aware dynamic transfer)입니다.</strong> 정적 배치만으로는 부족합니다. 실제 추론 중에는 배치 크기, 컨텍스트 길이, 동시 요청 수에 따라 부하가 시시각각 변합니다. ATSInfer는 런타임 상황을 보고 특정 텐서를 RAM에서 GPU로 승격하거나 반대로 강등합니다. 정적 배치가 “출발선”이라면, 동적 전송은 “주행 중 차선 변경”에 해당합니다.</p>

<p><strong>셋째, 비동기 CPU-GPU 조율(asynchronous coordination)입니다.</strong> CPU 연산, GPU 연산, 그리고 둘을 잇는 PCIe 전송을 서로 겹쳐(overlap) 실행합니다. 순진하게 구현하면 GPU가 CPU의 계산이나 데이터 전송을 기다리며 노는 시간이 생기는데, 이 조율 계층이 그 유휴 시간을 메웁니다. 논문은 이 덕분에 GPU SM(스트리밍 멀티프로세서) 평균 활용도가 약 70% 올라갔다고 보고합니다.</p>

<h2 id="논문이-보고한-실험-결과">논문이 보고한 실험 결과</h2>

<p>다시 강조하지만, 아래 수치는 <strong>논문이 보고한 값</strong>이며 저희가 직접 재현한 것이 아닙니다. ATSInfer는 아직 코드가 공개되지 않았고(트윗에서도 “연구자들이 llama.cpp 팀에 코드를 공유해 주길”이라는 반응이 있었습니다), 120B급 모델과 RTX 4090은 이 글을 쓰는 샌드박스에서 재현하기 어렵습니다. 그래서 저희는 재현 대신 <strong>구조 분석과 함의 정리</strong>에 집중합니다.</p>

<p>논문의 헤드라인 수치는 다음과 같습니다. 기존 하이브리드 시스템(llama.cpp의 레이어 단위 오프로딩 포함) 대비, prefill(첫 토큰까지의 처리량)은 최대 1.94배, decode(초당 생성 토큰)는 최대 3.29배 빨라졌습니다.</p>

<p><img src="/assets/images/atsinfer-hybrid-cpu-gpu-tensor-scheduling-results.png" alt="ATSInfer가 논문에서 보고한 최대 속도 향상 비교" /></p>

<p>실험 환경은 RTX 4090(24GB) 및 RTX 3060 시스템에 64GB RAM 구성이며, 검증에 사용한 모델은 다음과 같습니다.</p>

<ul>
  <li>Llama 3.1-70B (INT4)</li>
  <li>Qwen3-Next-80B-A3B (INT4)</li>
  <li>Qwen3.5-122B-A10B (INT4)</li>
  <li>GPT-OSS-120B (MXFP4)</li>
</ul>

<p>즉 24GB 한 장으로 122B 파라미터 모델(A10B, 활성 파라미터 기준으로는 더 작은 MoE)까지 구동했다는 것이 핵심 주장입니다. 여기서 두 가지를 분리해 읽어야 합니다. 첫째, “3.29배”는 특정 조건에서의 <strong>최댓값</strong>이지 모든 모델·모든 배치에서 나오는 평균이 아닙니다. 둘째, 이 이득은 근본적으로 “GPU에 안 들어가던 것을 넣어 돌린다”가 아니라 “어차피 CPU-GPU를 오갈 수밖에 없는 상황에서, 오가는 방식을 더 똑똑하게 만든다”에서 옵니다. 병목이 PCIe 대역폭이라는 물리 법칙은 그대로이므로, ATSInfer의 기여는 그 대역폭을 낭비 없이 쓰고 GPU 유휴 시간을 줄인 데 있습니다.</p>

<h2 id="thakicloud-제품-적용-시사점">ThakiCloud 제품 적용 시사점</h2>

<p>ThakiCloud의 <strong>ai-platform</strong>은 Kubernetes와 Kueue 기반으로 다양한 고객 환경에서 모델을 서빙하는 AI/ML 인프라입니다. ATSInfer 같은 텐서 단위 스케줄링은 저희가 특히 주목하는 흐름과 맞닿아 있습니다.</p>

<p>첫째, <strong>온프레미스·소버린 환경의 경제성</strong>입니다. 국내 공공·금융 고객처럼 데이터를 외부로 내보낼 수 없는 환경에서는 자체 GPU로 모델을 돌려야 합니다. 이때 H100 8장짜리 랙 대신 소비자용 GPU 몇 장으로 중대형 모델을 감당할 수 있다면, 초기 CAPEX가 극적으로 낮아집니다. ATSInfer의 실험이 보여 주는 것은, “VRAM이 모자라면 무조건 GPU를 더 사야 한다”는 전제가 텐서 배치 최적화로 상당 부분 완화될 수 있다는 점입니다. 물론 그 대가는 처리량 감소이므로, 저지연이 필수인 워크로드에는 부적합합니다. 이 트레이드오프를 워크로드별로 판단하는 것이 저희 서빙 계층의 역할입니다.</p>

<p>둘째, <strong>멀티테넌트 스케줄링과의 결합</strong>입니다. ATSInfer의 “로드 인식 동적 전송”은 단일 노드 안에서의 텐서 이동이지만, 그 발상은 클러스터 수준에서도 유효합니다. Kueue로 GPU 자원을 큐잉하고 할당할 때, 어떤 요청을 어떤 정밀도·어떤 오프로딩 프로파일로 처리할지를 부하에 따라 결정하는 정책은 저희가 이미 고민하는 영역입니다. 텐서 단위 프로파일링이 노드 안에서 자원 이득을 짜내듯, 클러스터 스케줄러는 노드 사이에서 같은 일을 합니다.</p>

<p>셋째, <strong>비용-품질 곡선의 재정의</strong>입니다. 저희는 <a href="/ko/llmops/ktransformers-moe-offload-28x-validation/">ktransformers 재현 글</a>에서 “28배” 같은 화제 수치가 숨은 전제 위에 서 있음을 직접 측정으로 보였습니다. ATSInfer의 “3.29배”도 같은 렌즈로 봐야 합니다. 마케팅 수치가 아니라, 우리 고객의 실제 모델·실제 배치·실제 SLA에서 어떤 숫자가 나오는지를 검증하는 것이 저희가 제공하는 가치입니다. 낮은 서빙 비용에서의 경쟁력은 결국 이런 검증의 축적에서 나옵니다.</p>

<h2 id="한계-및-반론">한계 및 반론</h2>

<p>가장 큰 한계는 <strong>코드 미공개</strong>입니다. 논문의 수치가 재현 가능한지, 다른 하드웨어·다른 모델에서도 유지되는지는 코드가 나와야 검증할 수 있습니다. 15,000줄 규모의 C++ 확장이라면 유지보수와 llama.cpp 본류 병합도 만만치 않은 과제입니다. 병합되지 못한 포크는 시간이 지나면 상류 변경을 따라가지 못해 쓸모가 줄어듭니다.</p>

<p>둘째, <strong>이득의 조건 의존성</strong>입니다. 텐서 단위 배치의 효과는 CPU 성능, RAM 대역폭, PCIe 세대에 크게 좌우됩니다. 논문의 실험은 64GB RAM을 전제하는데, RAM이 부족하면 텐서를 CPU에 둘 여유 자체가 없어집니다. PCIe 3.0 시스템에서는 전송이 병목이 되어 이득이 크게 줄어들 가능성이 높습니다([추정], 논문이 세대별 비교를 명시하지 않았습니다).</p>

<p>셋째, <strong>decode 최적화의 태생적 천장</strong>입니다. decode는 메모리 대역폭에 묶인(memory-bound) 작업입니다. 아무리 스케줄링을 잘해도 VRAM 밖에 있는 가중치는 매 토큰마다 어떤 식으로든 접근해야 하므로, 순수 VRAM 상주 대비 느릴 수밖에 없습니다. ATSInfer가 하는 일은 “느려지는 정도를 최소화”하는 것이지 “느려짐을 없애는” 것이 아닙니다. 반대편 논거를 세워 보면, 진짜 저지연·고처리량이 필요한 프로덕션 서빙이라면 여전히 모델 전체가 VRAM에 들어가는 GPU를 쓰는 편이 옳습니다. ATSInfer가 빛나는 지점은 “그 GPU를 살 여력이 없거나, 살 필요까지는 없는” 개발·평가·소규모 배치 구간입니다.</p>

<p>그럼에도 이 방향성은 분명한 가치가 있습니다. 하드웨어를 늘리지 않고 소프트웨어로 자원 활용을 짜내는 접근은, 온프렘·비용효율·self-hosting을 무기로 삼는 저희 같은 플랫폼에 특히 잘 맞습니다. 코드가 공개되면 저희 서빙 벤치마크에 편입해 실제 수치를 직접 측정할 계획입니다.</p>

<h2 id="출처">출처</h2>

<ul>
  <li>ATSInfer 논문: <a href="https://arxiv.org/abs/2607.10183">arXiv:2607.10183, Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices</a></li>
  <li>관련 글: <a href="/ko/llmops/ktransformers-moe-offload-28x-validation/">40만 달러 랙을 24GB로? ktransformers의 28배를 직접 재현해봤습니다</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="llmops" /><category term="ATSInfer" /><category term="llama.cpp" /><category term="CPU오프로딩" /><category term="GPU" /><category term="LLM서빙" /><category term="LLMOps" /><category term="양자화" /><category term="인프라" /><summary type="html"><![CDATA[레이어나 전문가 단위가 아니라 텐서 하나하나를 CPU와 GPU에 나눠 배치하면, RTX 4090 한 장으로 VRAM을 훌쩍 넘는 모델을 돌리면서 디코딩을 최대 3.29배 끌어올릴 수 있다는 논문 ATSInfer를 읽고, 그 원리와 우리 서빙 관점에서의 함의를 정리했습니다.]]></summary></entry><entry xml:lang="ko"><title type="html">카나리 게이트로 자율 스킬 배포의 안전판을 놓다: 트레이스 기반 회귀 탐지와 자동 롤백</title><link href="https://thakicloud.github.io/ko/research/canary-gated-skill-rollout/" rel="alternate" type="text/html" title="카나리 게이트로 자율 스킬 배포의 안전판을 놓다: 트레이스 기반 회귀 탐지와 자동 롤백" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/ko/research/canary-gated-skill-rollout</id><content type="html" xml:base="https://thakicloud.github.io/ko/research/canary-gated-skill-rollout/"><![CDATA[<p>이 글은 야간 무인 파이프라인이 스스로 만들어낸 스킬이나 설정을 다음 날 프로덕션에 어떻게 내보낼지 고민하는 국내 클라우드/AI 엔지니어를 위해 씁니다. 자율 하네스가 자기 코드를 고치는 사례가 늘어나는 지금, 그 고친 결과물을 어떤 절차로 실제 트래픽에 태울지는 아직 표준이 없습니다. 오늘 소개할 논문은 이 배포 단계에 카나리와 자동 롤백을 끼우는 구체적인 설계를 제안하고, 그 효과를 시뮬레이션으로 직접 재보았습니다.</p>

<p>문제의 출발점은 저자들이 직전에 내놓은 연구입니다. 1,600개가 넘는 스킬을 가진 프로덕션 하네스가 밤사이 스스로 라우팅과 검색 실패를 진단하고 리트리버 설정과 분해 규칙을 사람 검수 없이 재생성하는 구조였습니다. 문제는 배포 방식이었습니다. 재생성된 후보는 다음 날 아침 곧바로 전체 트래픽의 100퍼센트를 서빙했습니다. 오프라인 복구 루프 자체의 점검을 통과했더라도 실제 트래픽에서 더 나빠질 가능성은 남아 있었고, 그런 경우 사람이 알아챌 때까지 그날의 모든 요청이 회귀를 그대로 맞았습니다. 이 논문은 “생성된 후보”와 “전체 트래픽을 서빙하는 후보” 사이에 결정론적이고 트레이스로 촉발되는 롤백 게이트를 가진 점진적 트래픽 카나리를 끼우면, 이 무검증 전면배포 대비 장애 반경과 평균 복구시간(MTTR)이 얼마나 줄어드는지를 직접 잽니다.</p>

<p>첫 번째 기여는 게이트를 감으로 잡지 않았다는 점입니다. 슬라이딩 윈도우 회귀 탐지 게이트의 창 크기(W)와 오류 임계값(k)을 정확 이항분포 꼬리확률 탐색으로 골라, 기저 오류율 0.03에서 단일 윈도우의 오탐 확률이 2.12×10⁻⁶이 되도록 맞췄습니다. 저자들은 실험 과정도 숨기지 않았습니다. 초안에서는 창 크기 10에 임계 30퍼센트(오류 3개)라는 소박한 설정을 썼는데, 이 값이 기저 오류율에서도 오탐이 거의 확실하게 발생하는 수준이었다는 사실을 이항분포 캘리브레이션 과정에서 발견하고 최종 실험 전에 바로잡았다고 밝힙니다. 결과에서 발견한 게 아니라 결과를 내기 전에 계산으로 잡아낸 실수라는 점이 이 공개의 핵심입니다.</p>

<p>실험은 완전히 오프라인이고 시드가 고정된 재현 가능한 시뮬레이션입니다. 2만 틱짜리 스킬 호출 결과 스트림을 만들고 2천 틱째에 오류율이 계단식으로 튀는 회귀를 주입했습니다. 전면배포(카나리 가중치 1.0, v2가 매 틱마다 트래픽을 받음)와 카나리(가중치 0.1, 열 틱 중 한 번만 v2가 트래픽을 받음) 두 정책을 동일한 게이트 아래에서 비교했고, 회귀 심각도 세 단계(오류율 0.15, 0.35, 0.6, 기저값 0.03 대비)에 시드 60개씩을 곱해 총 360회를 돌렸습니다. 결과는 명확했습니다. 360회 전부에서 오탐은 0건, 탐지율은 100퍼센트였습니다.</p>

<p>카나리는 장애 반경(blast fraction)을 1.0에서 0.1로, 즉 90퍼센트 줄였습니다. 하지만 저자들은 이 숫자를 헤드라인으로 내세우지 않습니다. 카나리 가중치를 0.1로 정하고 탐지가 항상 성공한다면 장애 비율이 대략 0.1에 수렴하는 것은 설계상 거의 당연한 결과이기 때문입니다. 세 심각도 모두에서 장애 반경 감소율이 정확히 90.0퍼센트로 똑같이 나온 것 자체가, 이 숫자가 측정치가 아니라 라우팅 가중치의 구조적 산물이라는 증거라고 논문은 스스로 지적합니다.</p>

<p><img src="/assets/images/posts/research/canary-gated-skill-rollout/blast-fraction-by-severity.png" alt="Blast Fraction by Rollout Policy and Regression Severity" />
<em>카나리(가중치 0.1)는 전면배포(가중치 1.0) 대비 세 심각도 모두에서 장애 반경을 0.1로 낮춘다. 다만 논문은 이 90퍼센트 감소를 실증적 발견이 아니라 카나리 가중치 설계에 따른 구조적 결과로 명시한다. 시뮬레이션 데이터를 시각화한 분석용 그림이다.</em></p>

<p>측정치로서 의미가 더 큰 것은 v2에 실제로 노출된 절대 호출 건수입니다. 전면배포에서는 회귀가 발생한 뒤 평균 2,012에서 2,160건의 요청이 버그 있는 버전을 그대로 맞았습니다. 카나리에서는 그 수가 212에서 378건으로 줄었습니다. 심각도별로 82.5퍼센트(오류율 0.15)에서 89.5퍼센트(오류율 0.6)까지 감소했는데, 감소폭이 가장 작은 쪽이 오히려 완만한 회귀라는 점이 눈에 띕니다. 카나리는 완만한 회귀를 탐지하는 데 시간이 더 걸리고, 그만큼 v2 노출이 더 쌓이기 때문입니다.</p>

<p><img src="/assets/images/posts/research/canary-gated-skill-rollout/v2-exposed-invocations.png" alt="v2-Exposed Invocation Count by Policy and Severity" />
<em>전면배포는 약 2,000건의 호출을 버그 버전에 그대로 노출시키지만 카나리는 212에서 378건으로 줄인다. 감소폭이 가장 작은 쪽은 완만한 회귀로, 탐지 시간이 길어질수록 노출이 함께 늘어난다는 것을 보여준다. 저장소 자체의 로컬 벤치(2026년 7월 19일 측정)에서 나온 결과다.</em></p>

<p>이 논문이 실제로 “발견”이라 부를 만한 대목은 따로 있습니다. 카나리는 모든 심각도에서 전면배포보다 롤백까지 더 오래 걸렸지만, 그 격차는 회귀가 심할수록 급격히 좁혀졌습니다. 완만한 회귀(오류율 0.15)에서는 전면배포가 30,864.6밀리초 만에 롤백한 반면 카나리는 53,841.7밀리초가 걸려 약 74.5퍼센트 더 느렸습니다. 중간 심각도(오류율 0.35)에서는 그 격차가 8.86퍼센트로, 심한 회귀(오류율 0.6)에서는 4.82퍼센트로 줄었습니다. 원인은 게이트의 작동 방식에 있습니다. 게이트는 최소 30개의 v2 호출 결과를 모으고 그중 8개 이상이 오류여야 발동합니다. 전면배포는 매 틱이 v2 호출이라 이 30개를 거의 즉시 채우지만, 카나리는 열 틱 중 한 틱만 v2로 가기 때문에 같은 30개를 채우려면 전체 시스템 틱이 약 열 배 더 필요합니다. 완만한 회귀는 애초에 창을 여러 번 검사해야 임계값을 넘기 때문에 여기에 카나리의 트래픽 희석이 곱해져 지연이 크게 벌어집니다. 반대로 심한 회귀는 회귀가 시작된 직후 첫 창 검사에서 바로 임계값을 넘기 때문에, 두 정책의 차이는 그 첫 창 하나를 채우는 시간 차이로 줄어듭니다.</p>

<p><img src="/assets/images/posts/research/canary-gated-skill-rollout/mttr-penalty-by-severity.png" alt="Canary MTTR Penalty Relative to Full Rollout by Severity" />
<em>카나리의 MTTR 페널티는 완만한 회귀에서 약 74.5퍼센트에 달하지만 심한 회귀에서는 9퍼센트 아래로 줄어든다. 이 심각도별 격차가 논문이 실제로 새롭게 측정한 결과다. MTTR 수치는 이 저장소 자체의 관측성 스팬 로그에서 뽑은 실측 지연시간 6개(12.705, 35.029, 25.011, 6.271, 6.458, 0.004밀리초)를 시드 샘플링해 합산한 값이다.</em></p>

<p>한 가지 더 눈여겨볼 지점은 이 MTTR 수치가 상상으로 만든 숫자가 아니라는 것입니다. 저장소 안 관측성 파이프라인(<code class="language-plaintext highlighter-rouge">scripts/harness/span.py</code>의 스모크 테스트)이 실제로 남긴 스팬 로그 6개를 지연시간 풀로 삼아 시드 고정 샘플링했습니다. 표본이 여섯 개뿐이라는 한계는 저자들도 곧바로 인정합니다.</p>

<p>이 연구가 갖는 실용적 의미는 세 층위로 나뉩니다. ThakiCloud AI 플랫폼 입장에서는 밤샘 자율 스킬 진화와 회고 기반 모델 승격이 실제 스케줄 러너에 반영될 때, 카나리와 자동 롤백 게이트를 끼워 무검증 전면배포의 장애 반경을 줄이면 24시간 무인 운영의 안전 마진을 확보할 수 있다는 실용적 근거가 됩니다. 조금 더 넓게 보면, 에이전트가 스스로 코드나 스킬을 고치는 자기수정 AI가 늘어나는 흐름에서 결정론적 트레이스 기반 회귀 게이트와 자동 롤백을 묶는 이 패턴은 자율 시스템 안전성을 위한 재사용 가능한 표준 컴포넌트로 자리잡을 여지가 있습니다. 학술적으로는, 기존 카나리 배포 연구가 대체로 마이크로서비스나 웹 트래픽을 배포 단위로 다뤄온 데 비해 이 논문은 배포 단위를 “자율 에이전트가 생성하고 수정한 스킬”로 확장했습니다. LLM 산출물이 갖는 비결정론적 품질을, 모델의 자기보고가 아니라 스팬 트레이스 위의 결정론적이고 코드가 소유하는 게이트로 다룬다는 방법론이 자기진화 에이전트 하네스 문헌에 새 축을 더한다는 것입니다. 관련 연구에서 논문은 MOSS의 이진 헬스프로브와 전체 트래픽 스왑 방식을, 자신의 점진적 트래픽 카나리와 통계적 회귀 게이트 방식과 정확히 한 축에서 구분해 자리매김합니다.</p>

<p>논문은 스스로의 한계도 길게, 그리고 솔직하게 적어 둡니다. 가장 무거운 한계는 실험 전체가 완전히 합성된 결과 스트림이라는 점입니다. 실제 스킬이 실제 트래픽 분할에 배포된 적은 없고, 실제 롤백도 실행되지 않았습니다. 지연시간 풀이 단 6개의 실측 스팬으로 이루어져 실제 프로덕션 지연시간 분포나 꼬리 행동을 대표하지 못한다는 점, 90퍼센트 장애 반경 감소가 발견이 아니라 카나리 가중치 0.1과 완벽한 탐지율이 만든 구조적 결과라는 점도 반복해서 짚습니다. 게이트 설정(창 크기 30, 임계 8) 역시 단 하나만 시험했기 때문에 MTTR과 장애 반경 사이의 파레토 프론티어는 아직 그려지지 않았습니다. 회귀 모델도 계단형으로 상승하는 베르누이 오류율 하나뿐이라, 점진적 드리프트나 지연시간만 나빠지는 경우, 비용 폭증, 의미적 품질 저하 같은 다른 실패 양상에는 이 결과가 그대로 옮겨지지 않을 수 있습니다. 무엇보다 이 논문은 카나리와 롤백 메커니즘을 제안하고 시뮬레이션으로 검증했을 뿐, 실제 프로덕션 복구 루프나 하네스의 라이브 관측 코드에 직접 구현하지는 않았습니다.</p>

<p>더 자세한 내용과 전체 데이터는 Hugging Face 논문 페이지에서 확인할 수 있습니다.</p>

<p><a href="https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-20-canary-gated-skill-rollout">https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-20-canary-gated-skill-rollout</a></p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="research" /><category term="canary deployment" /><category term="agent harness" /><category term="skill evolution" /><category term="automatic rollback" /><category term="span tracing" /><category term="regression detection" /><category term="self-improving pipelines" /><category term="observability" /><category term="cost-latency monitoring" /><category term="autonomous agents" /><summary type="html"><![CDATA[밤새 스스로 코드를 고치는 에이전트 하네스가 다음 날 그 결과물을 어떻게 배포해야 안전할까요. 카나리와 자동 롤백을 끼운 실험이 장애 반경과 복구시간의 트레이드오프를 구체적으로 보여줍니다.]]></summary></entry></feed>