<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://thakicloud.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://thakicloud.github.io/" rel="alternate" type="text/html" /><updated>2026-07-20T12:32:14+09:00</updated><id>https://thakicloud.github.io/feed.xml</id><title type="html">Thaki Cloud Tech Blog | ThakiCloud | 다키클라우드 기술 블로그</title><subtitle>Thaki Cloud (ThakiCloud, 다키클라우드, thaki cloud, THAKI CLOUD, ثاكي كلاود)는 AI/ML Engineering, LLMOps, DevOps 분야의 최신 기술과 실무 경험을 공유하는 전문 기술 블로그입니다. 머신러닝 모델 운영, 쿠버네티스, 클라우드 인프라, AI 엔지니어링 커리어, 인공지능 기술 블로그, 다키클라우드 개발 팀의 깊이 있는 인사이트를 제공합니다. مدونة تقنية متخصصة في هندسة الذكاء الاصطناعي والحوسبة السحابية.</subtitle><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><entry xml:lang="ar"><title type="html">تشريح Kimi Code CLI: كيف يستحوذ وكيل الطرفية مفتوح المصدر على المحررات عبر ACP</title><link href="https://thakicloud.github.io/ar/agentops/kimi-code-cli-acp-open-source-agent/" rel="alternate" type="text/html" title="تشريح Kimi Code CLI: كيف يستحوذ وكيل الطرفية مفتوح المصدر على المحررات عبر ACP" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/ar/agentops/kimi-code-cli-acp-open-source-agent</id><content type="html" xml:base="https://thakicloud.github.io/ar/agentops/kimi-code-cli-acp-open-source-agent/"><![CDATA[<p>في الأسبوع الماضي، تصدرت Moonshot AI قوائم الترتيب في البرمجة بعد إطلاق نموذجها المفتوح الأوزان Kimi K3. لكن ما رافق هذا الإطلاق بهدوء كان أداة أقرب إلى سير عمل المطورين من النموذج نفسه، وهي Kimi Code CLI، وكيل برمجة طرفي مفتوح المصدر أطلقته Moonshot برخصة MIT. تداول مستخدمو LinkedIn تعريفاً يقول إن هذه الأداة تقدم ميزات غير موجودة في Claude Code. لم ننقل هذه العبارة كما هي، بل تحققنا مباشرة من المستودع الرسمي والوثائق. والخلاصة أن نصف هذا التعريف صحيح ونصفه الآخر مبالغ فيه. والنقطة الأكثر إثارة للاهتمام كانت في موضع لم تبرزه مواد الترويج.</p>

<p>يستعرض هذا المقال ماهية Kimi Code CLI، وما تقدمه فعلياً، ولماذا تستحق المتابعة من منظورنا كمشغّلين لمنصة ذكاء اصطناعي قائمة على K8s. وقد خصصنا مساحة واسعة لشرح لماذا يُعد معيار Agent Client Protocol المفتوح قطعة قد تغير قواعد اللعبة في منظومة الوكلاء.</p>

<h2 id="ما-هو-kimi-code-cli">ما هو Kimi Code CLI</h2>

<p>Kimi Code CLI أداة برمجة عاملة بأسلوب الوكيل تعمل من الطرفية، وتنتمي إلى نفس فئة Claude Code وGemini CLI وCodex CLI. المستودع الرسمي هو <a href="https://github.com/MoonshotAI/kimi-code">MoonshotAI/kimi-code</a>، وقد تطور من المشروع السابق <a href="https://github.com/MoonshotAI/kimi-cli">MoonshotAI/kimi-cli</a> مع الحفاظ على استمرارية الجلسات والإعدادات القديمة. كلا المستودعين رسميان من Moonshot، وينبغي الحذر من الخلط بينهما وبين مشاريع طرف ثالث تحمل أسماء مشابهة.</p>

<p>وهنا يجب توضيح تصحيح أول. الاسم الرسمي لهذه الأداة ليس “CLI مخصصة لـ Kimi K3” بل Kimi Code CLI. فهي أداة غير مقيدة بنموذج واحد، وتستخدم افتراضياً نموذج Moonshot المتخصص في البرمجة Kimi K2.7 Code، لكن يمكن عبر الإعدادات التحول إلى نماذج أخرى بما فيها K3. أي أن K3 هو أحد النماذج المتعددة التي يمكن ربطها بهذه الأداة، وليست الأداة مصممة حصراً له. أما K3 نفسه فهو نموذج MoE مفتوح بحجم 2.8 تريليون معلمة أطلقته Moonshot في 16 يوليو 2026، ويعتمد على Kimi Delta Attention مع نافذة سياق تصل إلى مليون رمز. وقد غطت هذا الإطلاق وسائل إعلام رئيسية مثل CNBC وBloomberg وForbes.</p>

<p>من المفيد رسم الصورة الكاملة أولاً لتترابط التفاصيل لاحقاً بسهولة أكبر. جوهر الأمر هو الدور المزدوج الذي تلعبه Kimi Code CLI: فمن جهة تتصل بالأدوات والبيانات كعميل MCP، ومن جهة أخرى تتصل بالمحررات كخادم ACP.</p>

<pre><code class="language-mermaid">flowchart TB
    subgraph EDITOR["محرر المطور (عميل ACP)"]
        ZED["Zed"]
        JB["عائلة JetBrains"]
        VSC["VS Code / Neovim"]
    end
    ACP["Agent Client Protocol&lt;br/&gt;JSON-RPC over stdio"]
    subgraph CLI["Kimi Code CLI (نواة الوكيل)"]
        MAIN["الوكيل الرئيسي&lt;br/&gt;يحافظ على سجل المحادثة"]
        SUB["الوكلاء الفرعيون&lt;br/&gt;coder · explore · plan&lt;br/&gt;لكل منهم سياق معزول"]
    end
    MODEL["طبقة النموذج&lt;br/&gt;Kimi K2.7 Code / K3&lt;br/&gt;أو نقطة نهاية متوافقة مع OpenAI"]
    subgraph MCP["خوادم MCP (الأدوات والبيانات)"]
        T1["Context7"]
        T2["Chrome DevTools"]
        T3["موصلات داخلية"]
    end

    EDITOR --&gt; ACP
    ACP --&gt;|kimi acp| MAIN
    MAIN --&gt; SUB
    MAIN --&gt;|طلب استدلال| MODEL
    SUB --&gt;|طلب استدلال| MODEL
    MAIN --&gt;|استدعاء أداة| MCP
</code></pre>

<h2 id="الوكلاء-الفرعيون-تقسيم-السياق-للحفاظ-على-نظافة-الوكيل-الرئيسي">الوكلاء الفرعيون: تقسيم السياق للحفاظ على نظافة الوكيل الرئيسي</h2>

<p>توفر Kimi Code CLI ثلاثة أنواع من الوكلاء الفرعيين المدمجين. الوكيل coder هو المسؤول الهندسي العام الذي يقرأ الملفات ويكتبها وينفذ الأوامر لتطبيق التغييرات الفعلية. الوكيل explore مخصص للاستكشاف، إذ يتصفح قاعدة الشيفرة للقراءة فقط. أما الوكيل plan فيقتصر عمله على تقديم خطط التنفيذ وتصاميم البنية دون تنفيذ أي أوامر في الصدفة. هذا التقسيم موثق في الصفحة الرسمية <a href="https://moonshotai.github.io/kimi-code/en/customization/agents.html">Agents and Sub-Agents</a>.</p>

<p>الجوهر هنا ليس الأسماء بل عزل السياق. يمتلك كل وكيل فرعي نافذة سياق مستقلة تماماً، ولا يرى سوى وصف المهمة الذي يمرره الوكيل الرئيسي صراحة. لا يُكشف سجل محادثة الوكيل الرئيسي للوكلاء الفرعيين، كما أن سجلات الاستدلال الوسيط واستدعاءات الأدوات التي ينفذها الوكيل الفرعي لا تختلط بسجل الوكيل الرئيسي، إذ يعيد الوكيل الفرعي النتيجة النهائية فقط. ولهذا السبب يظل السياق الرئيسي رفيعاً ولا يتضخم بالسجلات في الجلسات الطويلة. كما تدعم الأداة التنفيذ في الخلفية والتنفيذ المتوازي، بحيث يمكن تشغيل عدة مهام استكشاف في آن واحد وتعود النتائج تلقائياً عند الانتهاء.</p>

<p>هذا النمط ليس غريباً علينا. فحتى إطار التنسيق الداخلي الذي يشغّل هذه المدونة يفوّض مهام الاستكشاف إلى وكلاء فرعيين منخفضي التكلفة، ولا يستعيد سوى الملخصات لحماية السياق الرئيسي. مبدأ أن نظافة السياق تعني في الوقت نفسه التكلفة والجودة يبقى واحداً بغض النظر عن الأداة المستخدمة.</p>

<h2 id="mcp-تجربة-إعداد-دون-تعديل-json-يدوياً">MCP: تجربة إعداد دون تعديل JSON يدوياً</h2>

<p>تُدار عمليات ربط Model Context Protocol عبر مسارين. الأول هو الأوامر الفرعية لسطر الأوامر، حيث تُدار الخوادم عبر kimi mcp add وkimi mcp list وkimi mcp remove وkimi mcp authorize. على سبيل المثال يمكن ربط خادم بحث في الوثائق عبر نقل HTTP، أو ربط خادم أتمتة متصفح عبر نقل stdio.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># نقل HTTP (يدعم خيار OAuth)</span>
kimi mcp add <span class="nt">--transport</span> http context7 https://mcp.context7.com/mcp

<span class="c"># ربط عملية محلية عبر نقل stdio</span>
kimi mcp add <span class="nt">--transport</span> stdio chrome-devtools <span class="nt">--</span> npx chrome-devtools-mcp@latest
</code></pre></div></div>

<p>أما المسار الثاني فهو الأمر التفاعلي بشرطة مائلة /mcp-config الذي يُستخدم داخل واجهة TUI، ويتيح إضافة الخوادم وتعديلها والمصادقة عليها دون تحرير ملف إعدادات JSON مباشرة. ويعرض الأمر /mcp قائمة الخوادم المتصلة حالياً والأدوات المحمّلة. والجزء الذي أبرزه تعريف LinkedIn، وهو أنه لا حاجة لتعديل JSON مباشرة، صحيح فعلاً. لكن هذه الميزة بحد ذاتها ليست غائبة عن Claude Code، وسنعود إلى هذه النقطة لاحقاً. الوثائق ذات الصلة موجودة في <a href="https://moonshotai.github.io/kimi-cli/en/customization/mcp.html">إعداد MCP</a>.</p>

<h2 id="agent-client-protocol-أهم-قطعة-في-هذه-الأداة">Agent Client Protocol: أهم قطعة في هذه الأداة</h2>

<p>هذا هو الجزء الأكثر إثارة للاهتمام في هذا المقال. Agent Client Protocol، ويُختصر بـ ACP، هو معيار مفتوح صممه فريق محرر Zed. يعمل برخصة Apache، ويتبادل الرسائل عبر JSON-RPC 2.0 فوق stdio. تُشغّل المحررات الوكيل كعملية فرعية وتتواصل معه عبر المدخلات والمخرجات القياسية، وآلية النقل نفسها مطابقة لبروتوكول خادم اللغة.</p>

<p>التشبيه هنا يساعد كثيراً على الفهم. قبل ظهور LSP، كان على كل محرر أن يبني تكاملاً منفصلاً لكل لغة برمجة. حوّل LSP هذه المسألة من مشكلة M ضرب N إلى مشكلة M زائد N، فما إن ينفذ محرر واحد المعيار حتى يستفيد من أي خادم لغة بغض النظر عمن صنعه. ويفعل ACP الشيء ذاته تماماً مع الوكلاء، فما إن ينفذ محرر واحد ACP حتى يتصل به أي وكيل بطريقة موحدة بغض النظر عن صانعه. يمكن الاطلاع على هذا المفهوم في <a href="https://zed.dev/acp">تعريف Zed بـ ACP</a> وفي <a href="https://blog.marcnuri.com/agent-client-protocol-acp-introduction">مقال مارك نوري التوضيحي</a>.</p>

<p>من السهل الخلط بينه وبين MCP، لكن الاتجاه معاكس تماماً. يتجه MCP من الوكيل نحو الأدوات والبيانات، وفي هذه الحالة يكون الوكيل عميل MCP. أما ACP فيتجه من المحرر نحو الوكيل، وهنا يكون الوكيل خادم ACP والمحرر عميل ACP. أي أن الوكيل نفسه يلعب دور عميل MCP من جهة، ودور خادم ACP من جهة أخرى في آن واحد. وهذا هو السبب الذي جعل الرسم البياني السابق يوضح هذا الدور المزدوج.</p>

<p>تدعم Kimi Code CLI هذا البروتوكول بشكل أصلي عبر الأمر الفرعي kimi acp دون الحاجة إلى أي تثبيت إضافي. يتصل بها Zed بشكل أصلي، بينما تتصل به JetBrains عبر إضافة، وقد ظهرت بالفعل عدة تكاملات مع محررات أخرى وفق سجل ACP الخاص بـ Zed. بذلك يستطيع المطور تشغيل جلسة Kimi دون مغادرة المحرر الذي اعتاد عليه.</p>

<h2 id="إدخال-الصور-والفيديو-إلى-أي-حد-فعلياً">إدخال الصور والفيديو، إلى أي حد فعلياً</h2>

<p>ذكر تعريف LinkedIn أنه يمكن تمرير لقطة الشاشة كما هي كمدخل. وهذا يحتاج إلى تصحيح. فالميزة التي تبرزها Moonshot فعلياً ليست لقطة شاشة ثابتة بل إدخال مقطع فيديو مسجّل للشاشة. يذكر وصف المستودع أنه عند إسقاط تسجيل شاشة أو مقطع عرض توضيحي في المحادثة، يستطيع الوكيل مشاهدة وفهم السلوك الذي يصعب شرحه بالكلام مباشرة. وبالطبع تدعم نافذة الإدخال في سطر الأوامر لصق الصور أيضاً، إذ إن النموذج الافتراضي Kimi K2.7 Code هو نموذج متعدد الوسائط أصلي مزوّد بمُرمّز رؤية MoonViT بحجم 400 مليون معلمة، يستقبل النص والصور والفيديو معاً. غير أنه عند ربط نموذج مخصص، يجب تحديد دعم الصور صراحة ضمن modalities الخاصة بذلك النموذج ليعمل بشكل صحيح. وخلاصة القول إن إدخال الصور متاح فعلاً، لكن ما يُروَّج له كفارق حقيقي هو إدخال الفيديو، وتعبير لقطة شاشة غير دقيق تماماً.</p>

<h2 id="التثبيت-فعلياً-في-ثلاث-خطوات">التثبيت فعلياً في ثلاث خطوات</h2>

<p>تدفق التثبيت بسيط فعلاً كما ورد في التعريف. الأوامر أدناه مستندة إلى <a href="https://moonshotai.github.io/kimi-cli/en/guides/getting-started.html">دليل البدء الرسمي</a>، ولم نترك سجل تنفيذ مباشر لأن صندوق الاختبار الداخلي لدينا لا يملك صلاحية الوصول إلى نطاق التوزيع المعني. لذلك لم ننتج أي أرقام قياس أداء، واكتفينا بنقل الأوامر الموثقة فقط.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># 1) تشغيل سكربت التثبيت (يثبّت uv في الوقت نفسه)</span>
curl <span class="nt">-LsSf</span> https://code.kimi.com/install.sh | bash

<span class="c"># 2) التشغيل داخل دليل المشروع</span>
kimi

<span class="c"># 3) إعداد المصادقة</span>
/login
</code></pre></div></div>

<p>بالنسبة لنظام macOS تتوفر brew install kimi-code، وبالنسبة لويندوز يتوفر أيضاً سكربت PowerShell. للتطوير من الشيفرة المصدرية يلزم Node بإصدار 24.15 أو أحدث مع pnpm. ولأن الرخصة MIT، فإن القيود قليلة على قراءة الشيفرة وعمل fork لها وتوزيعها داخل المؤسسة.</p>

<h2 id="انفتاح-النماذج-ومزودي-الخدمة">انفتاح النماذج ومزودي الخدمة</h2>

<p>يبلغ أقصى طول للسياق في عائلة K2.6 نحو 256 ألف رمز، بينما يصل في K3 وفق مواد تسويق Moonshot إلى مليون رمز. لكن الأهم من ذلك هو انفتاح مزودي الخدمة. ففي ملف ~/.kimi-code/config.toml يمكن تسجيل عدة مزودين في آن واحد، من نقاط نهاية متوافقة مع OpenAI، إلى مفاتيح Anthropic API، وصولاً إلى Google GenAI أو Vertex AI. وهذا يعني أن الأداة غير مقيدة بنموذج واحد بعينه. كما تعالج تلقائياً حقل reasoning_content الخاص بنماذج الاستدلال من أطراف ثالثة. الوثائق ذات الصلة في <a href="https://moonshotai.github.io/kimi-cli/en/configuration/providers.html">Providers and models</a>.</p>

<h2 id="هل-هي-ميزات-غير-موجودة-في-claude-code-مقارنة-صريحة">هل هي ميزات غير موجودة في Claude Code: مقارنة صريحة</h2>

<p>أكثر العبارات انتشاراً في التعريف كانت أنها تقدم ميزات غير موجودة في Claude Code. وبعد التحقق تبين أن هذا الإطار في معظمه مبالغ فيه.</p>

<p>فالوكلاء الفرعيون وعزل السياق يقدمهما Claude Code أيضاً بالطريقة نفسها عبر ميزة الوكلاء الفرعيين. وMCP مدعوم أصلاً بنضج في Claude Code عبر نقل stdio وSSE وHTTP. كما أن لصق الصور موجود فيه أيضاً. هذه العناصر الثلاثة إذن ليست فوارق حقيقية.</p>

<p>الفارق الحقيقي يكمن في نقطتين. الأولى هي طريقة دعم ACP. فـ Kimi Code CLI تدمج ACP كميزة أساسية من الدرجة الأولى داخل الأداة نفسها عبر الأمر الفرعي kimi acp. أما Claude Code فيتصل بها عبر حزمة محول منفصلة صنعها Zed، وما تزال في مرحلة تجريبية. من منظور المستخدم، الأولى تعمل فور تفعيل الأداة، بينما الثانية تتطلب إضافة جسر إضافي. النقطة الثانية هي انفتاح النماذج. فـ Kimi مفتوح على سلسلة K مفتوحة الأوزان مع إمكانية التحول بين عدة مزودين، بينما يقتصر Claude Code على نماذج Anthropic حصرياً. ومن هذه النقطة يتفرع فارق ثالث يتعلق بإمكانية الاستضافة الذاتية. فبما أن Kimi أداة مفتوحة المصدر مع نموذج مفتوح الأوزان، يمكن تشغيله داخل المؤسسة، بينما تظل Claude Code أداة مفتوحة لكن نموذجها متاح عبر واجهة API فقط. يمكن الاطلاع على الدليل ذي الصلة في <a href="https://zed.dev/blog/claude-code-via-acp">مقال Zed حول Claude Code عبر ACP التجريبي</a>.</p>

<h2 id="دلالات-على-منتجات-thakicloud">دلالات على منتجات ThakiCloud</h2>

<p>يمس هذا الموضوع أداة وكيل من جهة، ومحور بنية تحتية يتعلق بالنماذج المفتوحة والاستضافة داخل المؤسسة من جهة أخرى. لذلك نستخدم العدستين معاً.</p>

<p>من عدسة Paxis، تتداخل بنية Kimi Code CLI إلى حد كبير مع اتجاه تصميم منتجنا. فPaxis هو مستوى التحكم الخاص بـ ThakiCloud لسحابة أصيلة الوكلاء (Agent-Native Cloud)، ويتعامل مع المهارات والأدوات والسياسات وسجلات التدقيق كموارد من الدرجة الأولى. والطريقة التي يعمل بها الوكلاء الفرعيون coder وexplore وplan لدى Kimi بالتوازي وفي سياقات معزولة، تشترك في الفلسفة نفسها مع طريقة عمل حاضنة المهارات في Paxis، التي تختار من بين أكثر من 960 مهارة باستخدام BM25 وتنفذها في صناديق اختبار معزولة. وACP بشكل خاص، بوصفه معياراً محايداً تجاه المزودين، يمثل فرصة مباشرة لـ Paxis. فأي وكيل ننشره، بما في ذلك الوكلاء المزودة بنماذج خضعت لضبط دقيق خاص بنا، يمكنه إذا نفذ ACP أن يتصل بمحررات المطورين لدى العملاء مثل Zed أو JetBrains بطريقة موحدة. وهذا المزيج من معيارين، MCP للاتصال بالبيانات وACP للاتصال بالمحررات، يمثل بالضبط الصورة التكاملية التي نتجه إليها.</p>

<p>ومن عدسة ai-platform، الانفتاح يعني حرية النشر مباشرة. فبوضع سلسلة K مفتوحة الأوزان فوق جدولة GPU عبر Kueue وخدمة vLLM في عنقودنا، وتوجيه الأداة نحو نقطة نهاية داخلية، يمكن بناء وكيل برمجة داخلي دون الاعتماد على API خارجي أو إخراج البيانات إلى الخارج. وهذا ينسجم مع متطلبات الأمن الخاصة بالاستضافة داخل المؤسسة في قطاعي المال والقطاع العام حيث لا يجوز خروج الشيفرة إلى الخارج، وكذلك مع متطلبات جهات مثل NIS. وقد سبق أن تناولنا في مقالات سابقة فكرة أنه كلما أصبحت القدرات شائعة ورخيصة، فإن ما تدفعه الشركات فعلياً هو بيئة تنفيذ محكومة. وأهمية Kimi Code CLI تكمن في أنها فتحت طبقة التنفيذ هذه كمصدر مفتوح.</p>

<h2 id="القيود-والاعتراضات">القيود والاعتراضات</h2>

<p>هناك عدة نقاط ينبغي النظر إليها بموضوعية. أولاً، أسماء المحركات الداخلية أو هياكل الطبقات التي تُذكر في تحليلات طرف ثالث معمقة لا ترد في الوثائق الرسمية، وقد تكون نتيجة هندسة عكسية، لذا من الأسلم الاستناد إلى الوثائق الرسمية قبل اعتبارها حقائق. ثانياً، توجد تقارير من المجتمع تفيد بأن مسار ACP يقدم جودة استجابة أفضل من طرق الاتصال الأخرى، لكنها انطباعات وليست قياسات أداء موثقة، أي أنها ليست أرقاماً محققة. ثالثاً، حتى مع كون النموذج مفتوح الأوزان، فإن خدمة نموذج بحجم 2.8 تريليون معلمة فعلياً داخل المؤسسة تتطلب موارد GPU كبيرة، والانفتاح لا يعني بالضرورة سهولة الاستضافة الذاتية، إذ يظل مسار API خياراً واقعياً للفرق الصغيرة. رابعاً، قد يتفوق نضج واستقرار منظومة الأدوات لدى Claude Code أو Codex CLI. وكون الأداة مفتوحة المصدر لا يعني بالضرورة أنها جاهزة للإنتاج.</p>

<p>ومع ذلك، فإن الاتجاه نحو ارتباط رخو بين الوكلاء والمحررات فوق معايير مفتوحة هو تيار واضح. فعالم يستطيع فيه المطور تبديل النموذج والمحرر كل على حدة دون التقيد بأداة سطر أوامر من مزود معين، هو عالم أكثر فائدة للمطورين. وتُعد Kimi Code CLI إحدى القطع التي تُقرّب هذا العالم.</p>

<h2 id="المصادر">المصادر</h2>

<ul>
  <li><a href="https://github.com/MoonshotAI/kimi-code">MoonshotAI/kimi-code (المستودع الرسمي)</a></li>
  <li><a href="https://github.com/MoonshotAI/kimi-cli">MoonshotAI/kimi-cli (المستودع السابق)</a></li>
  <li><a href="https://moonshotai.github.io/kimi-cli/en/guides/getting-started.html">دليل بدء Kimi Code CLI</a></li>
  <li><a href="https://moonshotai.github.io/kimi-code/en/customization/agents.html">وثائق Agents and Sub-Agents</a></li>
  <li><a href="https://moonshotai.github.io/kimi-cli/en/customization/mcp.html">وثائق إعداد MCP</a></li>
  <li><a href="https://zed.dev/acp">Zed - Agent Client Protocol</a></li>
  <li><a href="https://blog.marcnuri.com/agent-client-protocol-acp-introduction">ACP: The LSP for AI Coding Agents</a></li>
  <li><a href="https://zed.dev/blog/claude-code-via-acp">Zed - Claude Code via ACP (تجريبي)</a></li>
  <li><a href="https://www.marktechpost.com/2026/07/16/moonshot-ai-releases-kimi-k3-a-2-8-trillion-parameter-open-moe-model-with-kimi-delta-attention-and-1m-context/">MarkTechPost - إطلاق Kimi K3</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="agentops" /><category term="agentops" /><category term="kimi" /><category term="moonshot" /><category term="coding-agent" /><category term="mcp" /><category term="agent-client-protocol" /><category term="paxis" /><category term="thakicloud" /><summary type="html"><![CDATA[نحلل أداة سطر الأوامر المفتوحة المصدر للبرمجة التي أطلقتها Moonshot AI إلى جانب Kimi K3، استناداً إلى الوثائق الرسمية والمستودع الفعلي. من الوكلاء الفرعيين coder وexplore وplan، مروراً بإعداد MCP التفاعلي، وصولاً إلى الدعم الأصلي لـ Agent Client Protocol الذي يمثل الفارق الحقيقي، نتحقق إلى أي مدى تصح عبارة الترويج القائلة إن هذه ميزات غير موجودة في Claude Code.]]></summary></entry><entry xml:lang="ar"><title type="html">إعادة قراءة توقّعات ماسك في دافوس: الذكاء بعد خمس سنوات، وأين يعمل هذا الذكاء فعليًا</title><link href="https://thakicloud.github.io/ar/culture/three-year-window-labor-to-assets/" rel="alternate" type="text/html" title="إعادة قراءة توقّعات ماسك في دافوس: الذكاء بعد خمس سنوات، وأين يعمل هذا الذكاء فعليًا" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/ar/culture/three-year-window-labor-to-assets</id><content type="html" xml:base="https://thakicloud.github.io/ar/culture/three-year-window-labor-to-assets/"><![CDATA[<p><img src="/assets/images/three-year-window-labor-to-assets-hero.png" alt="صورة تجريدية لوقت الإنسان وهو يذوب في ضوء الروبوتات البشرية ومراكز البيانات" /></p>

<h2 id="نظرة-عامة">نظرة عامة</h2>

<p>انتشر مؤخرًا منشور في عدة مجتمعات. صيغ على هيئة رسالة إلى إيلون ماسك، وكان فحواه صارخًا. يقول إنه لم يبقَ أمام الناس سوى نحو ثلاث سنوات على الأكثر لكسب الرزق من العمل، وبعدها ستُؤتمت معظم الوظائف التي كانت تدفع راتبًا، ولن يبقى سبب لدفع المال لإنسان أصلًا. وسمّى المنشور ذلك أكبر انتقال للثروة في التاريخ، وخلص إلى أن المهمة الآن هي تحويل وقتك إلى أصول لا تستطيع الآلات انتزاعها منك.</p>

<p>لم يحمل المنشور نصًّا فحسب. فقد جاء معه مقطع فيديو حقيقي لماسك نفسه وهو يتحدث. لذلك سحبنا الفيديو المصدر وتحقّقنا، جملةً جملةً، مما قاله فعلًا. وكشف التحقّق أمرًا يستحق الملاحظة. فالرقم الذي انتشر، ثلاث سنوات، ليس ما قاله ماسك، والرقم الذي ذكره كان رقمًا مختلفًا. في هذا المقال نعرض أولًا ما قاله فعلًا، ثم ننظر في الصورة التي تتشكّل إذا أخذنا التوقّع على محمل الجد، وأخيرًا نسأل، من زاوية مدوّنة تقنية، أين يعمل هذا الذكاء فعليًا.</p>

<!-- Courtesy of embedresponsively.com -->

<div class="responsive-video-container">
    <iframe src="https://drive.google.com/file/d/1HczL43lXw-P-geWxtPEHQc_eiapkDPwx/preview" frameborder="0" webkitallowfullscreen="" mozallowfullscreen="" allowfullscreen=""></iframe>
  </div>

<p>المقطع أعلاه هو الذي انتشر. ويُذكر أنه جزء من حوار مع الرئيس التنفيذي لشركة BlackRock، لاري فينك، في المنتدى الاقتصادي العالمي في دافوس.</p>

<h2 id="ما-قاله-ماسك-فعلًا">ما قاله ماسك فعلًا</h2>

<p>عند تفريغه، تُختزل تصريحات ماسك في ثلاثة توقّعات. بكلماته: “خلال خمس سنوات، أي بحلول 2031 تقريبًا، أعتقد أن الذكاء الرقمي سيتجاوز مجموع كل الذكاء البشري.” “ستكون هناك، خلال خمس سنوات، ربما مئة مليون روبوت بشري على الأقل، وربما مليار.” “أتوقّع أن يبلغ حجم الاقتصاد ضعف حجمه الحالي خلال خمس، وربما ست أو سبع سنوات، لأنك ستدخل فترة مضاعفة، حيث يزداد الناتج الاقتصادي بسرعة كبيرة إلى حد أنك، مع فارق بضع سنوات زيادة أو نقصانًا، سترى تغيّرات هائلة.”</p>

<p>ثلاثة ادّعاءات إذًا. أولًا، خلال خمس سنوات تقريبًا يتجاوز الذكاء الرقمي مجموع كل الذكاء البشري. ثانيًا، خلال المدة نفسها تنمو الروبوتات البشرية إلى ما بين 100 مليون ومليار. ثالثًا، يتضاعف الاقتصاد خلال خمس إلى سبع سنوات. كل واحد من هذه على ساعة من خمس سنوات، لا ثلاث.</p>

<p>وهنا نقطة الدقّة التي يجدر تثبيتها. عبارة “ثلاث سنوات” التي انتشرت لم تكن تصريح ماسك، بل تأطيرًا أضافه صاحب المنشور فوق المقطع. تحدّث ماسك عن تغيّرات في الذكاء والروبوتات والاقتصاد على مدى خمس سنوات، فاختزل الكاتب الحافة الأمامية لذلك التغيّر، أي الوقت الذي يستطيع فيه المرء الصمود بالعمل، في رقم ثلاثة الدرامي. التمييز بين الرقمين هو حيث ينبغي أن يبدأ هذا النقاش. وفاءً للمصدر، الأفق الذي علينا مناقشته هو خمس سنوات تقريبًا، والموضوع هو العلاقة بين الذكاء والعمل، لا قائمة تسوّق من الأصول.</p>

<h2 id="لماذا-نأخذ-التوقّع-على-محمل-الجد">لماذا نأخذ التوقّع على محمل الجد</h2>

<p>تستحق التوقّعات المحدّدة زمنيًا الحذر، لكن اتجاه هذا التوقّع يقوم على أسس يصعب تجاهلها. نتّفق إلى حد كبير مع الاتجاه العام، ويجدر ذكر الأسباب مع الأدلّة.</p>

<p>أولًا، الذكاء. على مدى السنوات القليلة الماضية، ارتفعت قدرة نماذج اللغة على كتابة الشيفرة وصياغة الوثائق ومعالجة استفسارات العملاء وتحليل البيانات أسرع مما توقّع كثيرون. فالمهام التي ظل الناس يفترضون أن الآلات لا تستطيع أداءها سقطت واحدة تلو الأخرى. وطالما ظل الأداء يتحسّن على نحو يمكن التنبؤ به مع الحوسبة والبيانات المستثمرة في التدريب، فإن صورة ميل الذكاء الجمعي نحو الآلات في مرحلة ما ليست خيالًا بل امتدادًا للاتجاه.</p>

<p>ثانيًا، الروبوتات. الروبوتات البشرية هي ما يبدو عليه الأمر حين يتوقّف الذكاء عن العيش في البرمجيات فقط ويكتسب جسدًا يمتد به إلى العمل المادي. وسواء كان العدد 100 مليون أو مليارًا فذلك يتوقّف على الطاقة التصنيعية ومنحنيات الكلفة ويحمل هامش خطأ واسعًا، لكن الاتجاه نفسه، أي أن حصة معتبرة من العمل المادي تبدأ في الانتقال إلى الآلات، بات مرئيًا في المصانع وسلاسل الإمداد.</p>

<p>ثالثًا، الاقتصاد. حين يرتفع عرض الذكاء والعمل بحدّة يرتفع الناتج، وحين يرتفع الناتج بسرعة يدخل الاقتصاد فترة مضاعفة. توقّع ماسك بمضاعفة خلال خمس إلى سبع سنوات متفائل، لكن الآلية نفسها درسها علم الاقتصاد طويلًا. وحتى لو انزلق التوقيت بضع سنوات، يبقى الاتجاه قائمًا.</p>

<p>وبصدق، صحّة الاتجاه وصحّة التوقيت أمران مختلفان. تنطبق هنا ملاحظة روي أمارا القديمة على توقّعات التقنية: نميل إلى المبالغة في تقدير الأثر قصير المدى للتقنية والتقليل من أثرها بعيد المدى. قد يكون رقم الخمس سنوات خاطئًا. لكن الاتجاه أصعب في وصفه بالخطأ. من الأنزه أن نقول إن الاتجاه صحيح والسرعة غير مؤكّدة.</p>

<h2 id="من-العمل-إلى-الأصول-الجزء-المتين-والجزء-الهشّ">من العمل إلى الأصول: الجزء المتين والجزء الهشّ</h2>

<p>يبدأ استنتاج المنشور الأصلي بتعريف الراتب بأنه ثمن إعارتك لوقتك. فإذا استطاعت الآلات أداء عمل ذلك الوقت بكلفة أقل وثبات أكبر، تنخفض قيمة الوقت الذي تعيره، وفي النهاية لا يبقى إلا ما تملكه.</p>

<p>لنبدأ بالجزء المتين. إذا انقسم الدخل الذي ينتجه الاقتصاد بين العمل ورأس المال، فإن الأتمتة تميل بالكفّة نحو رأس المال. وأن يتدفّق العائد أكثر نحو من يملك الآلة كلما تولّت الآلات العمل البشري نمط تكرّر منذ الثورة الصناعية، وهذه الجولة من الذكاء الاصطناعي والروبوتات توسّعه عبر العمل المعرفي والعمل المادي معًا. والاتجاه، أي أن ما تملكه يصبح أكثر أهمية، ليس ادّعاءً غير معقول.</p>

<p>الآن الجزء الهشّ. أولًا، تاريخيًا تغيّرت الوظائف في شكلها أكثر مما اختفت كليًا. تزيل الأتمتة مهامًّا معيّنة بينما تخلق مهامًّا جديدة لتصميمها والإشراف عليها وربطها. لكن هذه المرة قد تُؤتمت المهام الجديدة نفسها بسرعة، فلا يمكن ببساطة تكرار التفاؤل القديم. ثانيًا، الوصفة القائلة “فاشترِ إذًا هذا الأصل بعينه الآن” تنطوي على قفزة. فثمة فجوة واسعة بين تشخيص ميل الدخل نحو رأس المال ووصفة شراء أصل معيّن اليوم. نحن لا نقدّم نصائح استثمارية، ونشير بوضوح إلى أن سعر أي أصل بعينه متقلّب. ما يتناوله هذا المقال ليس الرمز المتداول بل البنية الكامنة تحته.</p>

<h2 id="أين-يعمل-هذا-الذكاء">أين يعمل هذا الذكاء</h2>

<p>هنا نخطو خطوة من زاوية مدوّنة تقنية. الذكاء الرقمي الذي وصفه ماسك، وأدمغة الروبوتات البشرية، والناتج الذي يُفترض أنه يضاعف الاقتصاد، لا يحدث شيء من ذلك في الفراغ. فالحوسبة الفعلية التي تدرّب النماذج وتشغّل الاستدلال وتتحكّم في الروبوتات تحدث على وحدات GPU في مكان ما. بعبارة أخرى، الجوهر المادي لذكاء يتجاوز مجموع البشرية هو الحوسبة، والطاقة والبنية التحتية التي تُبقي تلك الحوسبة تعمل.</p>

<p>تنطبق هنا العبارة القديمة عن حمّى الذهب: المال الحقيقي ذهب لمن باع المعاول والسراويل أكثر مما ذهب لمن نقّب عن الذهب. وكلما جرت أتمتة الذكاء والعمل، ازدادت قيمة الحوسبة والطاقة اللتين تستهلكهما تلك الأتمتة. وهنا يظهر فارق حاسم عن أطروحة الأصول في المنشور الأصلي. فبعض الأصول تُملك لكنها لا تنتج شيئًا، بينما البنية التحتية للحوسبة تُملك وتنتج عملًا فعليًا في الوقت نفسه. إذا كان أحدهما وعاءً يخزّن القيمة، فالآخر مصنع يولّدها. في عصر الأتمتة، الجواب الأكثر إنتاجية عن سؤال “ماذا ينبغي أن أملك” هو أن تملك القدرة على تشغيل الأتمتة نفسها.</p>

<div class="mermaid">
flowchart TB
    A["وقت العمل البشري<br />(يُبادَل بالراتب)"] --&gt;|الاستبدال أو التعزيز بالأتمتة| B["عمل تؤدّيه الآلات<br />(عمل معرفي + عمل مادي)"]
    B --&gt; C{"إلى أين تتراكم<br />القيمة المُنتَجة؟"}
    C --&gt;|يخزّن القيمة| D["أصول نادرة<br />(وعاء لا ينتج شيئًا)"]
    C --&gt;|يولّد القيمة| E["الحوسبة والطاقة والبنية التحتية<br />(المصنع الذي يشغّل الذكاء)"]
    E --&gt; F["من يملك تلك البنية<br />ويتحكّم بها؟"]
    F --&gt; G["الأفراد: قدرة لا تُستبدَل<br />المؤسسات: حوسبتها وبياناتها الخاصة"]
</div>

<p>عند إنزالها إلى مستوى الفرد، تصبح هذه الرؤية أقل درامية لكنها أكثر عملية. فبدلًا من التسرّع في شراء الأصول، الاستعداد الأضمن لمعظم الناس هو أن تبني في داخلك القدرة على تصميم الأتمتة وتوجيهها والتحقّق منها، أي الحكم والسياق اللذين يصعب على الآلة استبدالهما. والانتقال من الخوف من الأداة إلى قيادتها هو الطريقة الأنزه لتحويل الوقت إلى أصل.</p>

<h2 id="من-منظور-thakicloud">من منظور ThakiCloud</h2>

<p>ثمة سبب يجعلنا لا نتعامل مع هذا الموضوع كأنه قصة شخص آخر. فما تبنيه ThakiCloud هو بالضبط تلك الأرضية التي يعمل عليها الذكاء.</p>

<p>العدسة الأولى هي ai-platform. إن ai-platform من ThakiCloud بنية تحتية للذكاء الاصطناعي والتعلّم الآلي تخصّص موارد GPU وتدرّب النماذج وتقدّم الاستدلال فوق Kubernetes. وثمة نقطة تحمل وزنًا خاصًا لدينا. فبدلًا من تسليم الحوسبة بالجملة لسحابة شخص آخر، نتيح للمؤسسة تشغيل نماذجها الخاصة على بياناتها الخاصة في بيئتها الخاصة. On-prem والذكاء الاصطناعي السيادي وكلفة التقديم المنخفضة والاستضافة الذاتية كلمات مفتاحية نعود إليها مرارًا. وبتطبيق المنطق أعلاه، الأمر يتعلّق بإعادة القدرة على امتلاك الحوسبة التي يستهلكها الذكاء والتحكّم بها إلى المؤسسات. ويثقل هذا المنظور أكثر في القطاع العام والصناعات المنظَّمة حيث تهمّ سيادة البيانات.</p>

<p>العدسة الثانية هي Paxis. امتلاك الحوسبة نصف الطريق فقط. فلا بدّ من مستوى تحكّم يحوّل تلك الحوسبة إلى عمل مؤتمت فعلي. Paxis سحابة أصلية للوكلاء تعمل فوق ai-platform، تعامل المهارات والأدوات والسياسات وسجلّات التدقيق كموارد من الدرجة الأولى. تختار من بين مئات المهارات لتشغيلها في صناديق رمل معزولة، وتنسج عدة وكلاء في مخطط DAG للتعاون، وتمرّر كل فعل عبر بوابة سياسات وسجلّ تدقيق. إذا كانت الحوسبة هي المصنع الذي يولّد القيمة، فإن Paxis أمر العمل لذلك المصنع وآلية أمانه. التقديم منخفض الكلفة يصنع اقتصاديات الأتمتة، وطبقة الوكلاء فوقه تحوّل تلك الاقتصاديات إلى نتائج عمل فعلية.</p>

<p>بجمع العدستين نحصل على جوابنا الخاص عن القلق الذي يثيره توقّع ماسك. فإذا كانت الأتمتة اتجاهًا لا مفرّ منه، فإن الحيلولة دون تركّز التحكّم في هذا التيار في أيدٍ قليلة وإتاحته لمزيد من المؤسسات هو، في رأينا، الدور الذي تستطيع شركة بنية تحتية أن تؤدّيه.</p>

<h2 id="ما-ينبغي-مراقبته-بحذر">ما ينبغي مراقبته بحذر</h2>

<p>مع الاتّفاق على الاتجاه العام، تبقى بضع نقاط ينبغي التمسّك بها حتى لا ننجرف.</p>

<p>أولًا، التوقيت. قد ينزلق رقم الخمس سنوات، ورقم الثلاث المنتشر أكثر. وإذ نتذكّر أن القيادة الذاتية الكاملة وُصفت بأنها “جاهزة العام المقبل” كل عام لأكثر من عقد، فكلما ثبّتت النبوءة تاريخًا بدقة أكبر، كان هدفها الشعور الذي تثيره أكثر من التاريخ نفسه. صدّق الاتجاه، لكن لا تصدّق الرزنامة.</p>

<p>ثانيًا، الوصفة. صحّة تشخيص ميل الدخل نحو رأس المال أمر مختلف عمّا ينبغي شراؤه الآن. فالأخير يُقدَّم عادة بنبرة يقين دون تحقّق. وكلما زادت الجملة يقينًا، ازدادت الحاجة إلى عادة التحقّق من أساسها على حدة.</p>

<p>أخيرًا، التوزيع. إذا كان انتقال الثروة كبيرًا فعلًا، فهو ليس مشكلة تُحلّ باختيارات الفرد الذكية للأصول بل مشكلة على المجتمع معالجتها معًا. والوصفة الموجّهة فقط لمن يقدرون على شراء الأصول تخاطر بتضخيم المشكلة نفسها التي شخّصتها. وقد طرح ماسك نفسه في مواضع أخرى مفاهيم مثل الدخل المرتفع الشامل، ما يشير إلى أن التوزيع بعد الأتمتة سؤال على مستوى المؤسسات لا الأفراد.</p>

<h2 id="خاتمة">خاتمة</h2>

<p>الشيء الوحيد الذي يمكننا أخذه بثقة من هذا المقطع هو التالي. الأتمتة حقيقية كاتجاه، وقيمتها تتراكم في القدرة على تشغيلها فعلًا: للأفراد، كقدرة لا تُستبدَل؛ وللمؤسسات، كحوسبتها وبياناتها الخاصة. وحتى لو لم تكن خمس سنوات ماسك رزنامة دقيقة، فإننا نتّفق إلى حد كبير مع الاتجاه القائل إن العلاقة بين الذكاء والعمل تتغيّر. غير أن ما وراء ذلك الباب ليس رمزًا متداولًا نتسرّع في شرائه، بل قدرة وبنية تحتية نبنيهما ببطء. هكذا نعيد قراءته.</p>

<h2 id="المصادر">المصادر</h2>

<ul>
  <li>الفيديو المصدر: مقطع لإيلون ماسك في حوار مع الرئيس التنفيذي لشركة BlackRock، لاري فينك، في المنتدى الاقتصادي العالمي في دافوس. اقتباسات هذا المقال فُرّغت مباشرة من ذلك الفيديو.</li>
  <li>تحقّق متقاطع من التصريحات: توقّعات ماسك الخمسية (تجاوز الذكاء الرقمي مجموع الذكاء البشري، ومئة مليون إلى مليار روبوت بشري، ومضاعفة الاقتصاد خلال خمس إلى سبع سنوات) نقلتها عدة وسائل إعلام بوصفها تصريحات دافوس.</li>
  <li>عبارة “ثلاث سنوات” المنتشرة هي تأطير صاحب المنشور الذي اقتبس المقطع، لا تصريح ماسك نفسه.</li>
  <li>ملاحظة روي أمارا (المبالغة قصيرة المدى والتقليل بعيد المدى) قاعدة إرشادية شائعة الاستشهاد في توقّعات التقنية.</li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="culture" /><category term="AIAutomation" /><category term="FutureOfWork" /><category term="ComputeInfrastructure" /><category term="SovereignAI" /><category term="AgentEconomy" /><category term="ElonMusk" /><category term="TechPhilosophy" /><summary type="html"><![CDATA[اختُزلت توقّعات إيلون ماسك الثلاثة في دافوس على الإنترنت إلى تحذير مفاده أن 'عصر الراتب ينتهي خلال ثلاث سنوات'. فرّغنا مقطع المقابلة الفعلي لنتحقق مما قاله حقًا، ثم ننظر في السبب الذي يجعل البنية التحتية للحوسبة التي تقوم تحت التحوّل من العمل إلى الأصول هي السؤال الحقيقي.]]></summary></entry><entry xml:lang="ar"><title type="html">فاتورة بقيمة 25 مليار وون تطرح سؤالا: لماذا تكلفة وكلاء الذكاء الاصطناعي غير مرئية؟</title><link href="https://thakicloud.github.io/ar/llmops/agent-cost-observability-billing-crisis/" rel="alternate" type="text/html" title="فاتورة بقيمة 25 مليار وون تطرح سؤالا: لماذا تكلفة وكلاء الذكاء الاصطناعي غير مرئية؟" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/ar/llmops/agent-cost-observability-billing-crisis</id><content type="html" xml:base="https://thakicloud.github.io/ar/llmops/agent-cost-observability-billing-crisis/"><![CDATA[<p>كُتب هذا المقال لمهندسي المنصات والبنية التحتية الذين يخططون لإدخال Claude Code أو وكلاء الذكاء الاصطناعي إلى مؤسساتهم، وللمسؤولين الماليين ومسؤولي المشتريات الذين سيضطرون لتفسير فاتورة الذكاء الاصطناعي في الشهر القادم. ولنبدأ بالخلاصة: أخبار تكلفة الذكاء الاصطناعي التي توالت خلال الشهر الأخير لا تدور حول أن “الذكاء الاصطناعي مكلف”. المشكلة الحقيقية هي أن <strong>الفاتورة لا تفسر ما الذي تمثله</strong>. ففي بنية الوكلاء، حيث يتحول طلب مستخدم واحد إلى عشرات أو مئات من استدعاءات النموذج وتنفيذ الأدوات، ثم إلى إعادة محاولة تلقائية عند الفشل، لا يمكن للمبلغ النهائي وحده أن يكشف أين تسربت الأموال داخل أي حلقة تنفيذ. نرى أن هذه الفجوة في المراقبة هي بالضبط جوهر الألم الذي يعيشه السوق الآن.</p>

<h2 id="نظرة-عامة">نظرة عامة</h2>

<p>يمكن تلخيص أخبار الفترة من أواخر يونيو حتى يوليو 2026 في سطر واحد: الاستخدام العشوائي للنماذج المتطورة (frontier models) في كل مهمة يجعل التكلفة يصعب تحملها، بل ويصعب حتى تتبع مصدرها. وقد ظهرت الحادثة بشكل درامي. فقد جرت محاولة تحصيل مبلغ يقارب 2.5 مليار وون من مستخدم محلي واحد، ثم محاولة تحصيل نحو 25 مليار وون لاحقا. ولحسن الحظ لم يُسحب أي مبلغ فعليا بسبب تجاوز حد البطاقة، غير أن تكرار وصول مبلغ غير طبيعي إلى مرحلة طلب موافقة البطاقة، وليس مجرد خطأ عرض، هو ما يمنح الحادثة وزنها المختلف.</p>

<p>وفي الفترة نفسها توالت أخبار من مستويات أخرى. فقد راجعت شركة تدقيق تكاليف الذكاء الاصطناعي فواتير 60 شركة وادّعت وجود فوترة زائدة كبيرة، وبدأت عدة شركات كبرى بتوزيع استخدام النماذج المتطورة على نماذج أرخص بحسب طبيعة المهمة، كما ورد أن شركات أمريكية وأوروبية انتقلت إلى نماذج صينية مفتوحة الأوزان بدافع خفض التكلفة. والمثير أن مزودي النماذج أنفسهم بدؤوا بالرد بما يفيد أن “تشغيل أفضل النماذج أداءً لفترات طويلة في كل مهمة أمر غير مستدام”. وكانت أخبار متعددة الاتجاهات تشير إلى النقطة نفسها.</p>

<h2 id="ماذا-حدث-في-الشهر-الأخير">ماذا حدث في الشهر الأخير</h2>

<p>أول ما لفت الانتباه كان مشكلة موثوقية نظام الفوترة. وبحسب تقرير ZDNet Korea، بدأ المبلغ المطلوب من طالب جامعي محلي بنحو 1.66 مليون دولار، ثم تضخم إلى نحو 16.62 مليون دولار أي عشرة أضعاف. وأوضحت أنثروبيك لاحقا أن الخطأ كان في إعداد مبلغ الشحن التلقائي الذي ضُبط بشكل غير طبيعي المرتفع، إلا أن المستخدم أكد أنه لم يُفعّل خاصية الشحن التلقائي إطلاقا، ما يترك سبب نشوء هذا الإعداد غامضا حتى الآن. وتُظهر واقعة إرسال المستخدم أكثر من خمس عشرة رسالة إلى عدة أقسام دون تلقي رد آلي إلا بعد مرور أربعة أيام، فجوةً في منظومة الاستجابة أوضح من كونها مجرد خطأ تقني.</p>

<p>المشكلة الثانية كانت في قابلية مراقبة تكلفة الوكلاء. وبحسب تقارير AI Times و The Information، أعلنت شركة ناشئة متخصصة في تدقيق التكاليف تُدعى Vaudit أنها راجعت فواتير 60 شركة بقيمة إجمالية نحو 34 مليون دولار، وخلصت إلى أن نحو 1.7 مليون دولار منها كانت فوترة زائدة. وشكّل استخدام Claude Code جزءا كبيرا من نطاق المراجعة، وذُكرت شركات مثل باناسونيك وHP وهوندا ضمن العملاء. وحسب ادعاء الشركة، فإن الأنماط شملت تسجيل استخدام نموذج رخيص بأسعار نموذج أغلى، وفرض رسوم على مهام لم تُنجز، وتكرار إعادة المحاولة التلقائية بعد الخطأ فيما يُعرف بـ<strong>عاصفة إعادة المحاولة (retry storm)</strong>. وهنا ينبغي توضيح نقطتين. أولا، ردت أنثروبيك بأنها لا تفرض رسوما على الطلبات غير المكتملة أو الاستجابات الخاطئة، وأنه لا يوجد دليل واسع على فوترة زائدة. ثانيا، تحصل Vaudit على نسبة من مبالغ الاسترداد الناجحة، أي أنها شركة تدقيق تجارية، لذا يجب قراءة هذه الأرقام كنتيجة تحقيق من طرف واحد لا كتدقيق حسابي مستقل. وبعبارة أخرى، نحن الآن في مرحلة تصادم بين ادعاءات شركة التدقيق ونفي المزود.</p>

<p>المشكلة الثالثة كانت في استجابة السوق. وذكرت The Information أن الشركات بدأت بفصل المهام: النماذج الرخيصة للتصنيف والتلخيص والتحويل البسيط، والنماذج المتطورة للبرمجة المعقدة ومهام الوكلاء، والنماذج مفتوحة الأوزان أو المستضافة ذاتيا للمهام المتكررة عالية الحجم. ونقلت الفايننشال تايمز أن شركات مثل DoorDash وSiemens وAirbnb اعتمدت نماذج DeepSeek أو من عائلة Moonshot لخفض التكلفة. وفي تقرير لـ Business Insider، اعترف حتى مسؤولو المنصة في أنثروبيك بأن ما يُعرف بـ<strong>التقنية المعلوماتية الظل (shadow IT)</strong>، أي اعتماد كل قسم لأدواته الخاصة بشكل منفصل، أدى إلى تضخم تكاليف الذكاء الاصطناعي في بعض الشركات، لكنهم شددوا على أن الحل ليس إيقاف الاستخدام أو فرض سقف ميزانية موحد، بل اختيار النموذج بحسب المهمة وإدارة مركزية للتكلفة على مستوى المؤسسة. كما تغيرت سياسات التسعير نفسها مرارا: تعدّلت مرات عدة شمولية النماذج عالية الأداء الأحدث ضمن الاشتراك وتوقيت التحول إلى الدفع بحسب الاستخدام، وتكرر تمديد مواعيد انتهاء العروض الترويجية. بل ورد أن إتاحة Claude Fable 5 مجانا امتدت حتى 19 يوليو. وكانت صعوبة توقع تكلفة الشهر القادم أكبر إزعاج لمسؤولي المشتريات، أكثر حتى من مسألة الأداء.</p>

<h2 id="لماذا-لا-يمكن-مراقبة-تكلفة-الوكلاء">لماذا لا يمكن مراقبة تكلفة الوكلاء</h2>

<p>السبب المشترك الذي يجمع بين هذه المسارات الثلاثة من الأخبار هو في النهاية واحد. ففي أحمال عمل الوكلاء، اتسعت المسافة كثيرا بين ما يراه المستخدم وما تسجله الفاتورة. كانت استدعاءات API التقليدية طلبا واحدا مقابل استجابة واحدة وسطر تكلفة واحد. أما وكلاء البرمجة أو حزم تطوير الوكلاء (agent SDK)، فإن أمرا واحدا فيها يتوسع إلى وضع خطة وتنفيذ واستدعاء أدوات وتحرير ملفات وتحقق، ثم إعادة محاولة عند الفشل. يحدث هذا التوسع في مكان لا يراه المستخدم، بينما تسجل الفاتورة مجموعه الكلي في سطر واحد فقط.</p>

<pre><code class="language-mermaid">flowchart TB
    U["طلب مستخدم واحد"] --&gt; P["وضع خطة الوكيل"]
    P --&gt; L["حلقة التنفيذ"]
    L --&gt; T["استدعاء أدوات · استدعاء نموذج&lt;br/&gt;عشرات إلى مئات المرات"]
    T --&gt; R{"هل نجحت؟"}
    R --&gt;|"فشل"| RS["إعادة محاولة تلقائية&lt;br/&gt;(retry storm)"]
    RS --&gt; T
    R --&gt;|"نجاح"| ACC["توكن · كاش · tool call&lt;br/&gt;تجميع تراكمي"]
    ACC --&gt; INV["الفاتورة: مبلغ نهائي في سطر واحد"]
    INV -.فجوة المراقبة.-&gt; U
</code></pre>

<p>في هذه البنية، تقع معظم نقاط تسرب التكلفة خارج مجال رؤية المستخدم. فحلقة إعادة المحاولة تعمل بصمت وتضخم عدد الاستدعاءات، ويحدث تباين بين الاستخدام الفعلي للنموذج والتفاصيل النهائية للفاتورة عند المرور عبر مزود سحابي وسيط، وإذا اختل إعداد واحد مثل الشحن التلقائي فقد يتدفق مبلغ غير طبيعي حتى مرحلة طلب موافقة البطاقة. تبدو الأخبار الثلاثة وكأنها حوادث منفصلة، لكنها في الواقع أوجه مختلفة لفجوة المراقبة نفسها. لذلك لا يكفي مجرد وضع حد شهري لكل مستخدم. المطلوب هو طبقة قياس مركزية تلتقط تكلفة كل نموذج، وتوكنات كل جلسة، وتوكنات الكاش، وعدد استدعاءات الأدوات، وتكلفة الفشل وإعادة المحاولة، ومعدل الزيادة اليومي غير الطبيعي <strong>في اللحظة التي يحدث فيها الاستدعاء نفسه</strong>. فبلا مراقبة لا توجد سيطرة، وبلا سيطرة تبقى الفاتورة دائما وثيقة مفاجأة بعد وقوع الحدث.</p>

<h2 id="دلالات-التطبيق-على-منتجات-thakicloud">دلالات التطبيق على منتجات ThakiCloud</h2>

<p>هذه المشكلة هي النقطة التي يستهدفها منتجا ThakiCloud من زاويتين مختلفتين. ولأن منظور البنية التحتية ومنظور الوكلاء يكملان بعضهما، نستخدم في هذا الموضوع العدستين معا.</p>

<p><strong>عدسة ai-platform: الملكية هي الحل لأحمال العمل المتكررة.</strong> الاستنتاج الذي وصل إليه السوق واضح. فمعالجة حتى المهام السهلة بنماذج متطورة يجعل التكلفة غير محتملة، بينما تصبح استضافة نموذج مفتوح الأوزان ذاتيا خيارا اقتصاديا للمهام المتكررة عالية الحجم. ومنصة ai-platform من ThakiCloud هي بالضبط بنية تحتية للذكاء الاصطناعي وتعلم الآلة قائمة على K8s مصممة لهذه النقطة. فهي تستخدم Kueue لجدولة وحدات GPU في طابور وزيادة معدل استخدامها، وتستخدم vLLM لخدمة النماذج مفتوحة الأوزان، وتعزل الاستخدام بحسب كل قسم عبر عزل متعدد المستأجرين مع فوترة منفصلة. فإذا كانت واجهات برمجة التطبيقات القائمة على الدفع بحسب الاستخدام تنتج فواتير يصعب توقعها، فإن الاستضافة الذاتية تبني بنية تكلفة GPU ثابتة لا يتقلب سعرها حتى مع تزايد الاستخدام. وعلى عكس الأنظمة الخارجية القائمة على الدفع بحسب الاستخدام التي تتغير سياساتها باستمرار، يحوّل النشر الداخلي أو السيادي إمكانية التنبؤ بالتكلفة نفسها إلى أصل قائم بذاته. كما أن عدم خروج البيانات إلى الخارج يمثل قيمة إضافية للمؤسسات ذات متطلبات التنظيم والأمن المحلية العالية.</p>

<p><strong>عدسة Paxis: جعل كل تصرف للوكيل قابلا للتدقيق.</strong> كان جوهر فجوة المراقبة هو حلقة الوكيل، وهذا هو بالضبط المجال الذي تعالجه Paxis. فـPaxis هي مستوى التحكم في السحابة الأصيلة للوكلاء (Agent-Native Cloud) من ThakiCloud، وتعمل فوق ai-platform، وتتعامل مع المهارات (Skills) والأدوات (Tools) والسياسات (Policies) وسجلات التدقيق (Audit Logs) كموارد من الدرجة الأولى. فكل ما يخص أي مهارة استدعاها الوكيل وبأي أداة وكم مرة، وفي أي بيئة معزولة تم التنفيذ، يُسجَّل بالكامل في سجل التدقيق. وفي هذه البنية، بدلا من أن تُضخّم عاصفة إعادة المحاولة الفاتورة بصمت، تظهر حلقة إعادة المحاولة بوضوح في سجل التدقيق، وتوقف بوابات السياسة الاستدعاءات التي تتجاوز العتبة المحددة. فتصميم يختار من بين أكثر من 960 مهارة عبر خوارزمية BM25 وينفذها في بيئة معزولة، ويُمرّر كل تصرف عبر السياسة والتدقيق، هو إجابة بنيوية بالضبط على مشكلة صعوبة معرفة أين نشأت التكلفة من مجرد النظر إلى الفاتورة. فالخدمة منخفضة التكلفة (ai-platform) تجعل الوكلاء اقتصاديين، بينما تجعل المراقبة على مستوى كل تصرف (Paxis) هذه الجدوى الاقتصادية قابلة للتنبؤ. وهكذا تتكامل العدستان.</p>

<h2 id="الحدود-والرأي-المعاكس">الحدود والرأي المعاكس</h2>

<p>من أجل التوازن، نوضح الجانب المعاكس بجلاء. أولا، ليست النماذج المتطورة تبذيرا بالضرورة. فبحسب تقرير وول ستريت جورنال، ترى شركات مثل Shopify أن النماذج المتطورة، في مهام البرمجة المعقدة والوكلاء متعددة الخطوات، توفر وقت المهندسين بما يبرر سعرها المرتفع. في المقابل، تتوخى شركات مثل Spotify وTwilio الحذر في تقييم ما إذا كان التحسن الطفيف في الأداء يبرر التكلفة الإضافية. أي أن الجواب ليس “تخلَّ عن النماذج المتطورة”، بل “وزّع بحسب صعوبة المهمة”. كما أن الاستضافة الذاتية ليست حلا شاملا لكل شيء؛ فتخفيض المهام التي تتطلب أعلى مستوى من الاستدلال إلى نماذج مفتوحة الأوزان يؤدي إلى تراجع الجودة، ويخلق عبئا تشغيليا جديدا يشمل تشغيل وحدات GPU وتحديث النماذج وتصحيحات الأمان.</p>

<p>ثانيا، الأرقام المتعلقة بالفوترة الزائدة المذكورة في هذا المقال ليست حقائق مؤكدة. فادعاء Vaudit هو إعلان من شركة تدقيق تجارية، وقد نفته أنثروبيك، لذا فإن الأدق حاليا هو قراءة الموقف على أنه تصادم بين طرفين. وحادثة الفوترة بمبلغ 25 مليار وون كذلك لم يُسحب فيها أي مبلغ فعليا، ولم يُكشف بعد عن تفسير تقني لسبب نشوء إعداد الشحن التلقائي. والاستنتاج الذي نخلص إليه من هذه الأخبار لا يستهدف مزودا بعينه، بل هو مبدأ مفاده أنه في عصر الوكلاء، وأيا كان المزود المستخدَم، يجب على الجهة المستخدِمة نفسها أن تؤمّن مراقبة التكلفة وحوكمتها. فمسألة اختيار نموذج جيد ومسألة ضبط ذلك النموذج مسألتان منفصلتان، وما كشفته أخبار الشهر الأخير هو أن المسألة الثانية كانت فارغة طوال الوقت.</p>

<h2 id="المصادر">المصادر</h2>

<ul>
  <li><a href="https://zdnet.co.kr/view/?no=20260709165452">ZDNet Korea، “مستخدم محلي تلقى طلب دفع بقيمة 25 مليار وون… جدل حول خطأ فوترة أنثروبيك” (2026-07-09)</a></li>
  <li><a href="https://zdnet.co.kr/view/?no=20260716093004">ZDNet Korea، “أنثروبيك التي طالبت بـ25 مليار وون: تبين أنه خطأ في إعداد الشحن التلقائي” (2026-07-16)</a></li>
  <li><a href="https://www.aitimes.com/news/articleView.html?idxno=212155">AI Times، “أنثروبيك في جدل ‘الفوترة الزائدة للذكاء الاصطناعي’: تحصيل رسوم حتى على المهام الفاشلة” (2026-06)</a></li>
  <li><a href="https://www.theinformation.com/titv/fedld">The Information، تقرير عن سيطرة الشركات على تكلفة الذكاء الاصطناعي وتوزيع النماذج (2026-06-23)</a></li>
  <li><a href="https://www.ft.com/content/9c8ff45b-7c20-4c2e-93c9-c52339ffdcee">Financial Times، “Companies turn to Chinese AI models to cut costs” (2026-07)</a></li>
  <li><a href="https://www.businessinsider.com/anthropic-ai-costs-responses-routers-2026-7">Business Insider، “Anthropic Official Warns Against ‘Wrong’ AI Cost Response” (2026-07-15)</a></li>
  <li><a href="https://www.wsj.com/cio-journal/meet-the-companies-shelling-out-for-top-ai-models-e1fe3375">The Wall Street Journal، “Meet the Companies Shelling Out for Top AI Models” (2026-07)</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="llmops" /><category term="LLMOps" /><category term="FinOps" /><category term="تكلفة الوكلاء" /><category term="مراقبة التكلفة" /><category term="توجيه النماذج" /><category term="self-hosting" /><category term="Paxis" /><category term="بنية الذكاء الاصطناعي التحتية" /><summary type="html"><![CDATA[من حادثة فوترة بقيمة 25 مليار وون طالت مستخدما محليا إلى شبهات فوترة زائدة شملت 60 شركة، تشير أخبار تكلفة الذكاء الاصطناعي في الشهر الأخير إلى فجوة واحدة. في عصر الوكلاء حيث يتحول الطلب الواحد إلى مئات من استدعاءات النموذج، لم تعد الفاتورة تفسر ما الذي دُفع المال مقابله. هذا المقال يلخص كيفية سد فجوة المراقبة هذه.]]></summary></entry><entry xml:lang="ar"><title type="html">نموذج بحجم 122B على بطاقة 24GB؟ شرّحنا ATSInfer الذي يقسّم llama.cpp على مستوى الموتّرات</title><link href="https://thakicloud.github.io/ar/llmops/atsinfer-hybrid-cpu-gpu-tensor-scheduling/" rel="alternate" type="text/html" title="نموذج بحجم 122B على بطاقة 24GB؟ شرّحنا ATSInfer الذي يقسّم llama.cpp على مستوى الموتّرات" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/ar/llmops/atsinfer-hybrid-cpu-gpu-tensor-scheduling</id><content type="html" xml:base="https://thakicloud.github.io/ar/llmops/atsinfer-hybrid-cpu-gpu-tensor-scheduling/"><![CDATA[<p>هذه المقالة موجّهة للمهندسين الذين يوازنون إمكانية تشغيل نموذج كبير ذاتياً على بطاقة GPU استهلاكية واحدة، ولمسؤولي البنية التحتية الذين يقرّرون مدى تصديق تغريدات “شغّل 120B على 24GB”. بدايةً: فكرة ATSInfer الأساسية (arXiv:2607.10183)، الصادرة عن باحثين في جامعة نانجينغ، بسيطة ومقنعة. حيث كان الإنزال (offloading) السابق ينقل الأشياء كتلاً على مستوى “الطبقة” أو “الخبير”، يقسّم ATSInfer وصولاً إلى <strong>الموتّرات (tensors) فرادى</strong>. ومع ذلك، فإن الرقم البارز “حتى 3.29 ضعفاً” يقوم على عدة افتراضات، والشيفرة لم تُنشر بعد. لم نُعِد إنتاج بطاقة RTX 4090 تشغّل نموذجاً بحجم 120B هنا، لذا فإن كل رقم في هذه المقالة هو <strong>قيمة أبلغت عنها الورقة</strong>، مذكورة على هذا الأساس.</p>

<h2 id="نظرة-عامة">نظرة عامة</h2>

<p>كل من شغّل نموذج LLM محلياً يصطدم بالجدار نفسه: إذا كانت أوزان النموذج أكبر من ذاكرة الـ GPU، يفيض الزائد إلى ذاكرة الـ CPU. وراية <code class="language-plaintext highlighter-rouge">-ngl</code> في llama.cpp (عدد الطبقات التي توضع على الـ GPU) تفعل هذا بالضبط. المشكلة أنها تقطع فقط على <strong>مستوى الطبقة</strong>. تخلط الطبقة الواحدة موتّرات مختلفة الطبيعة تماماً (أوزان attention، وأوزان FFN، ومعاملات التطبيع)، ومع ذلك تُعامَل بمبدأ الكل أو لا شيء: إما أن يذهب الكل إلى الـ GPU أو يبقى على الـ CPU.</p>

<p>سبب خسارة هذا الوضع الكتلي بسيط. قد يجعل نفس الـ 1GB في الـ VRAM موتّراً أسرع 10 أضعاف وآخر أسرع ضعفين فقط. الـ VRAM مورد نادر، والقطع الكتلي يمنعك من انتقاء الموتّرات ذات الربح الأعلى لكل GB. يستهدف ATSInfer هذا بالضبط. فهو يحلّل أداء كل موتّر على الـ CPU والـ GPU ويملأ الـ VRAM بدءاً من الموتّرات ذات <strong>أعلى ربح سرعة لكل GB</strong>. إذا كان ktransformers الأخير حيلةً على “مستوى الخبير” تدفع خبراء MoE إلى الـ CPU (انظر <a href="/ar/llmops/ktransformers-moe-offload-28x-validation/">مقالتنا ذات الصلة: إعادة إنتاج الـ 28 ضعفاً لـ ktransformers</a>)، فإن ATSInfer هو التعميم الأدق على “مستوى الموتّر”. واللافت أنه ينطبق على النماذج الكثيفة (dense) وليس MoE فقط.</p>

<h2 id="ما-هذه-التقنية">ما هذه التقنية</h2>

<p>ATSInfer نظام استدلال هجين CPU-GPU مبنيّ كامتداد لـ llama.cpp بنحو 15 ألف سطر من C++. كما يقول الاسم، فإن “الجدولة الآلية للموتّرات (Automated Tensor Scheduling)” هي الجوهر، وتتشابك ثلاث آليات.</p>

<pre><code class="language-mermaid">flowchart TB
    A["أوزان النموذج&lt;br/&gt;(RAM، تتجاوز سعة VRAM)"] --&gt; B["تحليل لكل موتّر&lt;br/&gt;قياس ربح السرعة لكل GB"]
    B --&gt; C{"التوزيع الثابت&lt;br/&gt;الموتّرات الأعلى ربحاً إلى VRAM أولاً"}
    C --&gt;|"موتّرات عالية الربح"| D["مقيمة في VRAM"]
    C --&gt;|"موتّرات منخفضة الربح"| E["مقيمة في RAM"]
    D --&gt; F["نقل ديناميكي واعٍ بالحمل&lt;br/&gt;ترقية/تنزيل حسب حمل التشغيل"]
    E --&gt; F
    F --&gt; G["تنسيق غير متزامن بين CPU وGPU&lt;br/&gt;تداخل الحساب ونقل PCIe"]
    G --&gt; H["إخراج الرموز&lt;br/&gt;prefill · decode"]
</code></pre>

<p><strong>أولاً، التوزيع الثابت للموتّرات (static placement).</strong> قبل تحميل النموذج، يقيس مدى تسارع كل موتّر على الـ GPU، ثم يضع الموتّرات على الـ GPU بترتيب “أكبر سرعة مُعادة لكل GB من الـ VRAM مستخدَم”. هذا قريب من مسألة حقيبة الظهر (knapsack) ويستغل مباشرةً التباين على مستوى الموتّر الذي تجاهله التوزيع الكتلي.</p>

<p><strong>ثانياً، النقل الديناميكي الواعي بالحمل (load-aware dynamic transfer).</strong> التوزيع الثابت وحده لا يكفي. خلال الاستدلال الفعلي، يتغير الحمل لحظةً بلحظة مع حجم الدفعة وطول السياق وعدد الطلبات المتزامنة. يرقّي ATSInfer موتّراً معيناً من الـ RAM إلى الـ GPU أو ينزّله بناءً على حالة التشغيل. إذا كان التوزيع الثابت هو خط الانطلاق، فالنقل الديناميكي هو تغيير المسار أثناء القيادة.</p>

<p><strong>ثالثاً، التنسيق غير المتزامن بين CPU وGPU (asynchronous coordination).</strong> يُداخِل بين حساب الـ CPU وحساب الـ GPU ونقل الـ PCIe الذي يصل بينهما. التنفيذ الساذج يترك الـ GPU خاملاً ينتظر عمل الـ CPU أو حركة البيانات؛ وطبقة التنسيق هذه تملأ ذلك الوقت الخامل. تُبلّغ الورقة أن هذا يرفع متوسط استغلال SM (المعالج المتعدد التدفق) في الـ GPU بنحو 70%.</p>

<h2 id="النتائج-التي-تبلّغ-عنها-الورقة">النتائج التي تبلّغ عنها الورقة</h2>

<p>مرة أخرى: الأرقام أدناه <strong>قيم أبلغت عنها الورقة</strong>، لا شيء أعدنا إنتاجه. شيفرة ATSInfer لم تُنشر بعد (حتى نقاش التغريدات قال “أيها الباحثون، رجاءً شاركوا الشيفرة مع فريق llama.cpp”)، وإعادة إنتاج نموذج بحجم 120B مع RTX 4090 صعبة في البيئة المعزولة وراء هذه المقالة. لذا بدلاً من إعادة الإنتاج، نركّز على <strong>التحليل البنيوي والدلالات</strong>.</p>

<p>عنوان الورقة العريض: مقابل الأنظمة الهجينة الحالية (بما فيها إنزال llama.cpp على مستوى الطبقة)، يتحسّن الـ prefill (الإنتاجية حتى أول رمز) حتى 1.94 ضعفاً، والـ decode (الرموز المولّدة في الثانية) حتى 3.29 ضعفاً.</p>

<p><img src="/assets/images/atsinfer-hybrid-cpu-gpu-tensor-scheduling-results.png" alt="أقصى تسارع تبلّغ عنه الورقة لـ ATSInfer" /></p>

<p>الإعداد هو نظام RTX 4090 (24GB) وRTX 3060 مع 64GB من الـ RAM، والنماذج المُتحقَّق منها هي:</p>

<ul>
  <li>Llama 3.1-70B (INT4)</li>
  <li>Qwen3-Next-80B-A3B (INT4)</li>
  <li>Qwen3.5-122B-A10B (INT4)</li>
  <li>GPT-OSS-120B (MXFP4)</li>
</ul>

<p>فالادّعاء المركزي هو تشغيل نموذج بـ 122 مليار معامل (وهو MoE بعدد معاملات نشطة أقل بكثير) على بطاقة 24GB واحدة. اقرأ هنا أمرين منفصلين. أولاً، “3.29 ضعفاً” هو <strong>حدّ أقصى</strong> تحت ظروف محددة، لا متوسط عبر كل نموذج وكل دفعة. ثانياً، الربح يأتي جوهرياً لا من “إدخال ما لم يكن يدخل في الـ GPU” بل من “جعل حركة CPU-GPU الحتمية أذكى”. تبقى الفيزياء نفسها في أن عرض نطاق الـ PCIe هو عنق الزجاجة، لذا فإن مساهمة ATSInfer هي استخدام ذلك العرض دون هدر وتقليل الوقت الخامل للـ GPU.</p>

<h2 id="دلالات-لمنتجات-thakicloud">دلالات لمنتجات ThakiCloud</h2>

<p>منصة <strong>ai-platform</strong> من ThakiCloud بنية تحتية للذكاء الاصطناعي/تعلّم الآلة تخدم النماذج عبر بيئات عملاء متنوعة على Kubernetes وKueue. الجدولة على مستوى الموتّر مثل ATSInfer تتوافق مع اتجاه نراقبه عن كثب.</p>

<p>أولاً، <strong>اقتصاديات البيئات المحلية (on-premises) والسيادية.</strong> في السياقات التي لا يمكن أن تغادر فيها البيانات المبنى، كعملاء القطاع العام والمالي المحليين، يجب أن تعمل النماذج على بطاقات GPU مملوكة. إذا استطاعت بضع بطاقات استهلاكية حمل نموذج متوسط إلى كبير بدلاً من رف من ثماني بطاقات H100، تنخفض النفقات الرأسمالية الأولية بشكل كبير. ما تُظهره تجارب ATSInfer هو أن الافتراض “إن نقص الـ VRAM فعليك ببساطة شراء المزيد من الـ GPU” يمكن تخفيفه جوهرياً عبر تحسين توزيع الموتّرات. الثمن، بالطبع، هو انخفاض الإنتاجية، لذا فهو غير مناسب لأحمال حرجة الكمون. الحكم على هذه المقايضة لكل حمل هو دور طبقة الخدمة لدينا.</p>

<p>ثانياً، <strong>الاقتران بالجدولة متعددة المستأجرين.</strong> “النقل الديناميكي الواعي بالحمل” في ATSInfer هو حركة موتّرات داخل عقدة واحدة، لكن الفكرة تصمد على مستوى العنقود أيضاً. عندما يصفّ Kueue موارد الـ GPU ويوزّعها، فإن سياسة تقرّر أي طلب يُعالَج بأي دقة وأي ملف إنزال بناءً على الحمل هي مجال نفكّر فيه بالفعل. تماماً كما يعصر التحليل على مستوى الموتّر مكاسب الموارد داخل العقدة، يفعل مجدول العنقود الشيء نفسه عبر العقد.</p>

<p>ثالثاً، <strong>إعادة تعريف منحنى الكلفة-الجودة.</strong> في <a href="/ar/llmops/ktransformers-moe-offload-28x-validation/">مقالة إعادة إنتاج ktransformers</a> أظهرنا بالقياس المباشر أن أرقاماً بارزة مثل “28 ضعفاً” تقوم على افتراضات خفية. و”3.29 ضعفاً” في ATSInfer يستحق العدسة نفسها. ليس رقماً تسويقياً، بل القيمة التي تظهر على نماذج عملائنا الفعلية ودفعاتهم الفعلية واتفاقيات مستوى الخدمة الفعلية، هي ما نتحقق منه. التنافسية عند كلفة خدمة منخفضة تأتي في النهاية من تراكم هذا النوع من التحقق بالضبط.</p>

<h2 id="الحدود-والحجج-المضادة">الحدود والحجج المضادة</h2>

<p>أكبر حدّ هو <strong>عدم نشر الشيفرة.</strong> لا يمكن التحقق مما إذا كانت أرقام الورقة قابلة لإعادة الإنتاج، وما إذا كانت تصمد على أجهزة ونماذج أخرى، إلا بعد ظهور الشيفرة. امتداد بـ 15 ألف سطر من C++ هو أيضاً مهمة صيانة ودمج مع المنبع (upstream) غير هيّنة. الفرع (fork) الذي لا يُدمج أبداً يتخلّف عن تغييرات المنبع مع الوقت ويفقد قيمته.</p>

<p>ثانياً، <strong>اعتماد المكاسب على الظروف.</strong> يعتمد أثر التوزيع على مستوى الموتّر بشدة على أداء الـ CPU، وعرض نطاق الـ RAM، وجيل الـ PCIe. تفترض تجارب الورقة 64GB من الـ RAM؛ فإن نقص الـ RAM لا يترك مجالاً لإبقاء الموتّرات على الـ CPU أصلاً. على أنظمة PCIe 3.0، يرجّح أن يصبح النقل عنق الزجاجة ويقلّص الربح جوهرياً ([تقدير]، لا تفصّل الورقة مقارنةً جيلاً بجيل).</p>

<p>ثالثاً، <strong>السقف المتأصّل لتحسين الـ decode.</strong> الـ decode مهمة مقيّدة بالذاكرة (memory-bound). مهما أحسنت الجدولة، يجب الوصول إلى الأوزان خارج الـ VRAM بطريقة ما مع كل رمز، فيكون حتماً أبطأ من الإقامة الكاملة في الـ VRAM. ما يفعله ATSInfer هو “تقليل مقدار التباطؤ”، لا “إزالة التباطؤ”. وبناءً على الحجة المعاكسة: لخدمة إنتاجية تحتاج حقاً كموناً منخفضاً وإنتاجية عالية، يبقى استخدام بطاقة GPU تحمل النموذج كاملاً في الـ VRAM هو القرار الصائب. يتألق ATSInfer في مجال التطوير والتقييم والدفعات الصغيرة حيث لا تقدر على تلك البطاقة أو لا تحتاجها.</p>

<p>ومع ذلك، للاتجاه قيمة واضحة. عصر استغلال الموارد عبر البرمجيات بدلاً من إضافة العتاد يناسب منصةً مثل منصتنا بشكل خاص، منصةً تتخذ من العمل المحلي وكفاءة الكلفة والاستضافة الذاتية أسلحةً لها. حالما تُنشر الشيفرة، نخطط لإدماجها في معايير الخدمة لدينا وقياس الأرقام الحقيقية بأنفسنا.</p>

<h2 id="المصادر">المصادر</h2>

<ul>
  <li>ورقة ATSInfer: <a href="https://arxiv.org/abs/2607.10183">arXiv:2607.10183, Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices</a></li>
  <li>مقالة ذات صلة: <a href="/ar/llmops/ktransformers-moe-offload-28x-validation/">رفّ بـ 400 ألف دولار على 24GB؟ أعدنا إنتاج الـ 28 ضعفاً لـ ktransformers</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="llmops" /><category term="ATSInfer" /><category term="llama.cpp" /><category term="CPU-offloading" /><category term="GPU" /><category term="LLM-serving" /><category term="LLMOps" /><category term="quantization" /><category term="infrastructure" /><summary type="html"><![CDATA[بدلاً من إنزال طبقات أو خبراء كاملين، يوزّع ATSInfer الموتّرات فرادى بين الـ CPU والـ GPU. تُبلّغ الورقة عن تشغيل نماذج تتجاوز سعة الـ VRAM بكثير على بطاقة RTX 4090 واحدة مع رفع سرعة الـ decode حتى 3.29 ضعفاً. قرأنا arXiv:2607.10183 وفكّكنا آلية العمل ومعناها للخدمة (serving).]]></summary></entry><entry xml:lang="en"><title type="html">Inside Kimi Code CLI: How an Open Source Terminal Agent Swallows Editors via ACP</title><link href="https://thakicloud.github.io/en/agentops/kimi-code-cli-acp-open-source-agent/" rel="alternate" type="text/html" title="Inside Kimi Code CLI: How an Open Source Terminal Agent Swallows Editors via ACP" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/en/agentops/kimi-code-cli-acp-open-source-agent</id><content type="html" xml:base="https://thakicloud.github.io/en/agentops/kimi-code-cli-acp-open-source-agent/"><![CDATA[<p>Last week Moonshot AI released the open-weight model Kimi K3 and took the top spot on the coding leaderboards. But something that touches developer workflows even more directly slipped out quietly alongside it: <strong>Kimi Code CLI</strong>, an open source terminal coding agent that Moonshot released under the MIT license. LinkedIn timelines were full of posts framing it as “features Claude Code doesn’t have.” We didn’t take that line at face value and instead checked the official repository and documentation ourselves. The short version: half of the pitch holds up, and half is overstated. The genuinely interesting part turned out to be something the marketing didn’t emphasize at all.</p>

<p>This post covers what Kimi Code CLI is, what it actually delivers, and why it is worth watching from the perspective of a team running a Kubernetes based AI platform. We spend a good part of it on why Agent Client Protocol, an open standard, is a piece that could reshape the agent ecosystem.</p>

<h2 id="what-kimi-code-cli-is">What Kimi Code CLI Is</h2>

<p>Kimi Code CLI is an agentic coding tool that runs in the terminal, in the same family as Claude Code, Gemini CLI, and Codex CLI. The official repository is <a href="https://github.com/MoonshotAI/kimi-code">MoonshotAI/kimi-code</a>, and it evolved from the earlier project <a href="https://github.com/MoonshotAI/kimi-cli">MoonshotAI/kimi-cli</a>, carrying existing sessions and configuration forward. Both repositories are official Moonshot projects, so be careful not to confuse them with similarly named third-party projects.</p>

<p>Here is the first correction worth making. The tool’s official name is not “the CLI for Kimi K3” but <strong>Kimi Code CLI</strong>. It is not tied to a single model. By default it pairs with Moonshot’s coding-specialized model, Kimi K2.7 Code, but configuration lets you switch to K3 or other models as well. K3 is one of several models the CLI can attach to, not a model the CLI was purpose-built for. K3 itself is a 2.8 trillion parameter open MoE model that Moonshot released on July 16, 2026, built around Kimi Delta Attention and up to a 1 million token context window. This launch was covered jointly by CNBC, Bloomberg, and Forbes, among other major outlets.</p>

<p>Establishing the full picture first makes the details land much better later. The core idea is that Kimi Code CLI plays a dual role: on one side it is an MCP client connecting to tools and data, and on the other side it is an ACP server connecting to editors.</p>

<pre><code class="language-mermaid">flowchart TB
    subgraph EDITOR["Developer Editor (ACP Client)"]
        ZED["Zed"]
        JB["JetBrains family"]
        VSC["VS Code / Neovim"]
    end
    ACP["Agent Client Protocol&lt;br/&gt;JSON-RPC over stdio"]
    subgraph CLI["Kimi Code CLI (Agent Core)"]
        MAIN["Main Agent&lt;br/&gt;Maintains conversation history"]
        SUB["Sub-agents&lt;br/&gt;coder · explore · plan&lt;br/&gt;Each with an isolated context"]
    end
    MODEL["Model Layer&lt;br/&gt;Kimi K2.7 Code / K3&lt;br/&gt;or OpenAI-compatible endpoint"]
    subgraph MCP["MCP Servers (Tools · Data)"]
        T1["Context7"]
        T2["Chrome DevTools"]
        T3["Internal connectors"]
    end

    EDITOR --&gt; ACP
    ACP --&gt;|kimi acp| MAIN
    MAIN --&gt; SUB
    MAIN --&gt;|inference request| MODEL
    SUB --&gt;|inference request| MODEL
    MAIN --&gt;|tool call| MCP
</code></pre>

<h2 id="sub-agents-splitting-context-to-keep-the-main-thread-clean">Sub-agents: Splitting Context to Keep the Main Thread Clean</h2>

<p>Kimi Code CLI ships with three built-in sub-agents. <strong>coder</strong> is the general engineering agent that reads and writes files and runs commands to make actual changes. <strong>explore</strong> is read-only and dedicated to surveying the codebase. <strong>plan</strong> produces implementation plans and architectural designs without running any shell commands. This split is spelled out in the official <a href="https://moonshotai.github.io/kimi-code/en/customization/agents.html">Agents and Sub-Agents</a> documentation.</p>

<p>What matters here is not the naming but the context isolation. Each sub-agent gets a fully independent context window and sees only the task description the main agent explicitly hands it. The main agent’s conversation history is never exposed to a sub-agent, and the intermediate reasoning and tool call logs a sub-agent produces never leak back into the main history. A sub-agent returns only its final conclusion. That is why the main context stays thin instead of bloating with logs over a long session. Background and parallel execution are also supported, so multiple exploration tasks can run at once and their results flow back automatically once complete.</p>

<p>This pattern is not unfamiliar to us. The internal orchestration harness behind this blog also delegates exploration to low-cost sub-agents and pulls back only summaries to protect the main context. The principle that context hygiene is both a cost and a quality lever holds regardless of which tool implements it.</p>

<h2 id="mcp-a-configuration-experience-without-hand-editing-json">MCP: A Configuration Experience Without Hand-Editing JSON</h2>

<p>Model Context Protocol integration is managed through two paths. The first is the CLI subcommand set: <code class="language-plaintext highlighter-rouge">kimi mcp add</code>, <code class="language-plaintext highlighter-rouge">kimi mcp list</code>, <code class="language-plaintext highlighter-rouge">kimi mcp remove</code>, and <code class="language-plaintext highlighter-rouge">kimi mcp authorize</code> manage servers. For example, you can attach a documentation search server over HTTP transport, or a browser automation server over stdio transport.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># HTTP transport (with optional OAuth)</span>
kimi mcp add <span class="nt">--transport</span> http context7 https://mcp.context7.com/mcp

<span class="c"># stdio transport connecting to a local process</span>
kimi mcp add <span class="nt">--transport</span> stdio chrome-devtools <span class="nt">--</span> npx chrome-devtools-mcp@latest
</code></pre></div></div>

<p>The second path is the interactive slash command <code class="language-plaintext highlighter-rouge">/mcp-config</code> inside the TUI, which lets you add, edit, and authenticate servers without touching a JSON config file directly. <code class="language-plaintext highlighter-rouge">/mcp</code> shows the currently connected servers and the list of loaded tools. The claim that “you don’t need to edit JSON directly,” which the LinkedIn post emphasized, is accurate. That said, this convenience by itself is not something Claude Code lacks either; we return to that point later in this post. See the <a href="https://moonshotai.github.io/kimi-cli/en/customization/mcp.html">MCP configuration</a> docs for details.</p>

<h2 id="agent-client-protocol-the-most-important-piece-of-this-tool">Agent Client Protocol: The Most Important Piece of This Tool</h2>

<p>This is the most interesting part of this article. Agent Client Protocol, abbreviated ACP, is an open standard built by the Zed editor team. It is Apache licensed and runs JSON-RPC 2.0 over stdio, with the editor spawning the agent as a child process and communicating through standard input and output. The transport mechanism itself is identical to the Language Server Protocol.</p>

<p>An analogy helps a great deal here. Before LSP existed, every editor needed a separate integration for every language. LSP turned that M-times-N problem into M-plus-N: an editor only needs to implement the standard once, and it benefits from every language server anyone builds. ACP does exactly the same thing for agents. Once an editor implements ACP, any agent built by anyone plugs into it through a standard interface. This concept is explained in <a href="https://zed.dev/acp">Zed’s introduction to ACP</a> and in <a href="https://blog.marcnuri.com/agent-client-protocol-acp-introduction">Marc Nuri’s write-up</a>.</p>

<p>It is easy to confuse ACP with MCP, but the direction is reversed. MCP flows from the agent toward tools and data, with the agent acting as the MCP client. ACP flows from the editor toward the agent, with the agent acting as the ACP server and the editor acting as the ACP client. The same agent can play the MCP client role on one side and the ACP server role on the other at the same time, which is exactly why the diagram above draws that dual role.</p>

<p>Kimi Code CLI supports this protocol natively through the <code class="language-plaintext highlighter-rouge">kimi acp</code> subcommand, with no separate installation required. Zed connects natively, JetBrains connects through a plugin, and based on Zed’s ACP registry, several other editor integrations are already listed. Developers can drive a Kimi session without ever leaving the editor they already use.</p>

<h2 id="image-and-video-input-what-actually-works">Image and Video Input: What Actually Works</h2>

<p>The LinkedIn post claimed you can “feed a screen capture directly as input.” This needs a correction. The feature Moonshot actually markets front and center is not a static screenshot but <strong>screen recording video input</strong>. The repository description says that dropping a screen recording or demo clip into the chat lets the agent directly see and understand behavior that’s hard to describe in words. Pasting images into the CLI input is also supported, and the default model, Kimi K2.7 Code, is natively multimodal with a 400 million parameter vision encoder called MoonViT, so it accepts text, images, and video all together. That said, when you attach a custom model, that model’s modalities must explicitly declare image support for this to work correctly. To summarize, image input does work, but the feature actually marketed as the real differentiator is video input, and the word “screenshot” is somewhat inaccurate here.</p>

<h2 id="installation-is-really-three-steps">Installation Is Really Three Steps</h2>

<p>The installation flow is as simple as advertised. The commands below follow the <a href="https://moonshotai.github.io/kimi-cli/en/guides/getting-started.html">official getting started guide</a>. Our internal sandbox does not have access to the relevant distribution domain, so we did not run these ourselves and are not logging execution output. We are not fabricating any benchmark numbers here; we are only citing verified commands.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># 1) Run the install script (also installs uv)</span>
curl <span class="nt">-LsSf</span> https://code.kimi.com/install.sh | bash

<span class="c"># 2) Run it from your project directory</span>
kimi

<span class="c"># 3) Set up authentication</span>
/login
</code></pre></div></div>

<p>macOS also supports <code class="language-plaintext highlighter-rouge">brew install kimi-code</code>, and Windows has a PowerShell script. Building from source requires Node 24.15 or later and pnpm. The license is MIT, which puts few restrictions on reading the code, forking it, or deploying it internally.</p>

<h2 id="model-and-provider-openness">Model and Provider Openness</h2>

<p>Context length tops out at 256,000 tokens for the K2.6 line, and Moonshot’s marketing claims up to 1 million tokens for K3. More important is provider openness. In <code class="language-plaintext highlighter-rouge">~/.kimi-code/config.toml</code> you can register multiple providers, including OpenAI-compatible endpoints, an Anthropic API key, and Google GenAI or Vertex AI. That means the CLI is not locked into a single model. It also automatically handles the <code class="language-plaintext highlighter-rouge">reasoning_content</code> field from third-party reasoning models. See <a href="https://moonshotai.github.io/kimi-cli/en/configuration/providers.html">Providers and models</a> for the details.</p>

<h2 id="is-this-a-feature-claude-code-lacks-an-honest-comparison">Is This a Feature Claude Code Lacks? An Honest Comparison</h2>

<p>The most widely circulated pitch was “this gives you features Claude Code doesn’t have.” After verifying it, that framing turns out to be mostly overstated.</p>

<p>Sub-agents and context isolation are already offered by Claude Code through its own sub-agent feature, in the same way. MCP is also mature in Claude Code, which already supports stdio, SSE, and HTTP transports. Image pasting exists in Claude Code as well. None of these three are actual differentiators.</p>

<p>The real differences sit in two places. First, how ACP support is delivered. Kimi Code CLI bakes ACP into the CLI itself as a first-class feature through the <code class="language-plaintext highlighter-rouge">kimi acp</code> subcommand. Claude Code, by contrast, connects through a separate adapter package built by Zed, currently in beta. From a user’s perspective, the former just works once you turn it on, while the latter requires bolting on an additional bridge. Second, model openness. Kimi is open across its open-weight K series with the ability to switch providers, while Claude Code is limited to Anthropic’s own models. A third difference follows from this: self-hosting potential. Kimi combines an open source CLI with open-weight models, making on-premise serving possible, whereas Claude Code has an open CLI but a model that is API-only. The relevant evidence is in <a href="https://zed.dev/blog/claude-code-via-acp">Zed’s post on the Claude Code ACP beta</a>.</p>

<h2 id="implications-for-thakiclouds-products">Implications for ThakiCloud’s Products</h2>

<p>This topic touches both the agent tooling axis and the open model and on-premise infrastructure axis, so we apply two lenses together.</p>

<p>Through the Paxis lens, Kimi Code CLI’s structure overlaps quite a bit with our product’s design direction. Paxis is ThakiCloud’s Agent-Native Cloud control plane, treating skills, tools, policies, and audit logs as first-class resources. The way Kimi’s coder, explore, and plan sub-agents run in parallel with isolated contexts shares the same underlying philosophy as the way Paxis’s skill harness selects among more than 960 skills with BM25 and runs them in isolated sandboxes. ACP in particular, as a vendor-neutral standard, is a direct opportunity for Paxis. Any agent we deploy, including ones running our own fine-tuned models, could plug into a customer’s development editor such as Zed or JetBrains through a standard interface simply by implementing ACP. The combination of MCP for connecting to data and ACP for connecting to editors is exactly the kind of integrated picture we are aiming for.</p>

<p>Through the ai-platform lens, openness translates directly into deployment freedom. Running an open-weight K-series model on top of our cluster’s Kueue GPU scheduling and vLLM serving, and routing the CLI to an internal endpoint, would let us build an internal coding agent without depending on an external API or exporting data outside our environment. That aligns well with on-premise security requirements in domains like finance or the public sector where code cannot leave the organization, including requirements tied to Korea’s National Intelligence Service. As capability becomes commoditized and cheaper, what enterprises actually end up paying for is a controlled execution environment, a point we have made in earlier posts as well. Kimi Code CLI matters precisely because it opens up that execution layer as open source.</p>

<h2 id="caveats-and-counterarguments">Caveats and Counterarguments</h2>

<p>A few things deserve a colder look. First, some third-party deep dives reference internal engine names or layer structures that don’t appear in the official documentation, which suggests they may be the product of reverse engineering. It is safer to treat the official docs as the source of truth before citing those as fact. Second, there are community reports that the ACP path produces better response quality than other connection methods, but this is anecdotal rather than benchmarked, and we don’t treat it as verified data. Third, even with open weights, actually serving a 2.8 trillion parameter class model on-premise requires substantial GPU resources. Openness does not automatically mean easy self-hosting, and the API route remains the practical choice for small teams. Fourth, the maturity and stability of the tooling ecosystem may still favor Claude Code or Codex CLI. Being open source does not by itself mean production ready.</p>

<p>Even so, the direction toward agents and editors loosely coupled through open standards is a clear trend. A world where developers aren’t locked into a single vendor’s CLI, and can swap out the model and the editor independently, is a better world for developers. Kimi Code CLI is one of the pieces bringing that world closer.</p>

<h2 id="sources">Sources</h2>

<ul>
  <li><a href="https://github.com/MoonshotAI/kimi-code">MoonshotAI/kimi-code (official repository)</a></li>
  <li><a href="https://github.com/MoonshotAI/kimi-cli">MoonshotAI/kimi-cli (previous repository)</a></li>
  <li><a href="https://moonshotai.github.io/kimi-cli/en/guides/getting-started.html">Kimi Code CLI getting started guide</a></li>
  <li><a href="https://moonshotai.github.io/kimi-code/en/customization/agents.html">Agents and Sub-Agents documentation</a></li>
  <li><a href="https://moonshotai.github.io/kimi-cli/en/customization/mcp.html">MCP configuration documentation</a></li>
  <li><a href="https://zed.dev/acp">Zed - Agent Client Protocol</a></li>
  <li><a href="https://blog.marcnuri.com/agent-client-protocol-acp-introduction">ACP: The LSP for AI Coding Agents</a></li>
  <li><a href="https://zed.dev/blog/claude-code-via-acp">Zed - Claude Code via ACP (beta)</a></li>
  <li><a href="https://www.marktechpost.com/2026/07/16/moonshot-ai-releases-kimi-k3-a-2-8-trillion-parameter-open-moe-model-with-kimi-delta-attention-and-1m-context/">MarkTechPost - Kimi K3 launch</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="agentops" /><category term="agentops" /><category term="kimi" /><category term="moonshot" /><category term="coding-agent" /><category term="mcp" /><category term="agent-client-protocol" /><category term="paxis" /><category term="thakicloud" /><summary type="html"><![CDATA[We dig into the open source coding CLI that Moonshot AI released alongside Kimi K3, working strictly from the official docs and repository. From coder/explore/plan sub-agents to interactive MCP configuration and its real differentiator, native Agent Client Protocol support, we verify how much of the 'features Claude Code lacks' marketing line actually holds up.]]></summary></entry><entry xml:lang="en"><title type="html">Rereading Musk’s Davos Forecast: Intelligence in Five Years, and Where That Intelligence Actually Runs</title><link href="https://thakicloud.github.io/en/culture/three-year-window-labor-to-assets/" rel="alternate" type="text/html" title="Rereading Musk’s Davos Forecast: Intelligence in Five Years, and Where That Intelligence Actually Runs" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/en/culture/three-year-window-labor-to-assets</id><content type="html" xml:base="https://thakicloud.github.io/en/culture/three-year-window-labor-to-assets/"><![CDATA[<p><img src="/assets/images/three-year-window-labor-to-assets-hero.png" alt="Abstract image of human time dissolving into the light of humanoids and data centers" /></p>

<h2 id="overview">Overview</h2>

<p>A post recently made the rounds across several communities. It was framed as a message to Elon Musk, and its thrust was stark. People have at most about three years left to earn a living through labor, it argued, after which most of the jobs that once paid a salary will be automated and there will no longer be a reason to pay a human at all. The post called this the largest wealth transfer in history and concluded that the task now is to convert your time into assets that machines cannot take from you.</p>

<p>The post did not carry only text. It came with an actual video clip of Musk himself speaking. So we pulled the source video and checked, sentence by sentence, what he actually said. The check turned up something worth noting. The number that went viral, three years, was not something Musk said, and the number he did give was a different one. In this piece we first lay out what he actually said, then look at the picture that emerges if we take the forecast seriously, and finally ask, from a tech-blog angle, where this intelligence actually runs.</p>

<!-- Courtesy of embedresponsively.com -->

<div class="responsive-video-container">
    <iframe src="https://drive.google.com/file/d/1HczL43lXw-P-geWxtPEHQc_eiapkDPwx/preview" frameborder="0" webkitallowfullscreen="" mozallowfullscreen="" allowfullscreen=""></iframe>
  </div>

<p>The clip above is the one that went viral. It is reported to be part of a conversation with BlackRock CEO Larry Fink at the World Economic Forum in Davos.</p>

<h2 id="what-musk-actually-said">What Musk Actually Said</h2>

<p>Transcribed, Musk’s remarks compress into three predictions. In his own words: “In five years, so five years being say 2031, I think digital intelligence will exceed the sum of all human intelligence.” “There will be, in five years, probably at least a hundred million humanoid robots, but maybe a billion.” “I will predict that the economy is probably twice its current size in five, maybe six, seven years, because you are going to hit a doubling period, where economic output is increasing so fast that, plus or minus a few years, you will see giant changes.”</p>

<p>Three claims, then. First, in about five years digital intelligence surpasses the sum of all human intelligence. Second, over the same span humanoid robots grow to somewhere between 100 million and 1 billion. Third, the economy doubles within five to seven years. Every one of these is on a five-year clock, not a three-year one.</p>

<p>Here is the accuracy point worth pinning down. The viral “three years” was not Musk’s statement but a framing the post’s author added on top of the clip. Musk spoke about changes to intelligence, robots, and the economy over five years, and the author compressed the front edge of that change, the time a person can hold on through labor, into a dramatic three. Distinguishing the two numbers is where this discussion has to begin. Faithful to the source, the horizon we should be discussing is around five years, and the subject is the relationship between intelligence and labor, not a shopping list of assets.</p>

<h2 id="why-we-take-the-forecast-seriously">Why We Take the Forecast Seriously</h2>

<p>Time-bound predictions deserve caution, but the direction of this one rests on grounds that are hard to dismiss. We largely agree with the broad direction, and it is worth stating the reasons alongside the evidence.</p>

<p>First, intelligence. Over the past few years, the ability of language models to write code, draft documents, handle customer inquiries, and analyze data has climbed faster than many expected. Tasks people kept assuming machines could not do have fallen one after another. As long as performance keeps improving predictably with the compute and data poured into training, the picture of collective intelligence tipping toward machines at some point is not fantasy but an extension of the trend.</p>

<p>Second, robots. Humanoids are what it looks like when intelligence stops living only in software and gains a body to extend into physical labor. Whether the count is 100 million or 1 billion depends on manufacturing capacity and cost curves and carries wide error bars, but the direction itself, that a meaningful share of physical work begins shifting to machines, is already visible on factory floors and in logistics.</p>

<p>Third, the economy. When the supply of intelligence and labor rises sharply, output rises, and when output rises fast, an economy enters a doubling period. Musk’s five-to-seven-year doubling is optimistic, but the mechanism is one economics has long studied. Even if the timing slips by a few years, the direction holds.</p>

<p>To be honest, the direction being right and the timing being right are two different things. Roy Amara’s old observation fits technology forecasts well: we tend to overestimate the short-run effect of a technology and underestimate the long-run one. The five-year figure may be off. But the direction is harder to call wrong. It is more honest to say the direction is right and the speed is uncertain.</p>

<h2 id="from-labor-to-assets-the-solid-part-and-the-soft-part">From Labor to Assets: The Solid Part and the Soft Part</h2>

<p>The original post’s conclusion starts by defining a salary as the price of lending out your own time. If machines can do the work of that time more cheaply and reliably, the value of the time you lend falls, and in the end only what you own remains.</p>

<p>Start with the solid part. If the income an economy produces splits between labor and capital, automation tilts the scale toward capital. That returns flow more to whoever owns the machine as machines take over human work is a pattern repeated since the Industrial Revolution, and this round of AI and robotics is widening it across both knowledge work and physical work. The direction, that what you own comes to matter more, is not an unreasonable claim.</p>

<p>Now the soft part. First, jobs have historically changed shape more often than they have vanished outright. Automation removes certain tasks while creating new ones to design, supervise, and stitch them together. This time, though, the newly created tasks may themselves be automated quickly, so the old optimism cannot simply be repeated. Second, the prescription to “therefore buy this particular asset now” contains a leap. There is a wide gap between the diagnosis that income is tilting toward capital and the prescription to buy a given asset today. We do not provide investment advice, and we note plainly that the price of any particular asset is volatile. What this piece deals with is not the ticker but the structure beneath it.</p>

<h2 id="where-that-intelligence-runs">Where That Intelligence Runs</h2>

<p>Here we step in from a tech-blog angle. The digital intelligence Musk described, the brains of humanoid robots, and the output that supposedly doubles the economy do not happen in thin air. The actual computation that trains models, runs inference, and controls robots happens on GPUs somewhere. In other words, the physical substance of an intelligence that exceeds the sum of humanity is compute, and the power and infrastructure that keep that compute running.</p>

<p>The old gold-rush line fits here: the real money went less to those who dug for gold than to those who sold the picks and jeans. The more intelligence and labor get automated, the more the compute and power that automation consumes are worth. And here a crucial difference from the original post’s asset thesis appears. Some assets are held but produce nothing, whereas compute infrastructure is held and produces actual work at the same time. If one is a vessel that stores value, the other is a factory that generates it. In an age of automation, the most productive answer to “what should I own” is to own the ability to run the automation itself.</p>

<div class="mermaid">
flowchart TB
    A["Human labor time<br />(exchanged for salary)"] --&gt;|replaced / augmented by automation| B["Work done by machines<br />(knowledge + physical labor)"]
    B --&gt; C{"Where does the<br />created value accrue?"}
    C --&gt;|stores value| D["Scarce assets<br />(a vessel that produces nothing)"]
    C --&gt;|generates value| E["Compute, power, infra<br />(the factory that runs intelligence)"]
    E --&gt; F["Who owns and<br />controls that infra?"]
    F --&gt; G["Individuals: irreplaceable capability<br />Organizations: their own compute and data"]
</div>

<p>Brought down to the individual, this insight is less dramatic but more practical. Rather than rushing to buy assets, for most people the surer preparation is to build within yourself the ability to design, direct, and verify automation, the judgment and context a machine cannot easily replace. Moving from fearing the tool to commanding it is the most honest way to turn time into an asset.</p>

<h2 id="through-thakiclouds-lens">Through ThakiCloud’s Lens</h2>

<p>There is a reason we do not treat this topic as someone else’s story. What ThakiCloud builds is precisely that floor on which intelligence runs.</p>

<p>The first lens is ai-platform. ThakiCloud’s ai-platform is AI/ML infrastructure that allocates GPU resources, trains models, and serves inference on top of Kubernetes. One point carries special weight for us. Instead of handing compute wholesale to someone else’s cloud, we let an organization run its own models on its own data in its own environment. On-prem and sovereign AI, low serving cost, and self-hosting are the keywords we return to again and again. Applied to the logic above, it is about returning to organizations the ability to own and control the compute that intelligence consumes. The angle weighs most in public-sector and regulated industries where data sovereignty matters.</p>

<p>The second lens is Paxis. Owning compute is only half. There has to be a control plane that turns that compute into actual automated work. Paxis is an Agent-Native Cloud running on top of ai-platform, treating skills, tools, policies, and audit logs as first-class resources. It selects among hundreds of skills to run in isolated sandboxes, weaves multiple agents into a DAG to collaborate, and passes every action through a policy gate and audit log. If compute is the factory that generates value, Paxis is that factory’s work order and its safety mechanism. Low-cost serving creates the economics of automation, and the agent layer above turns that economics into actual business outcomes.</p>

<p>Put the two lenses together and you get our own answer to the anxiety Musk’s forecast provokes. If automation is an unavoidable direction, then keeping control of that current from concentrating in a few hands and letting more organizations share it is, we believe, the part an infrastructure company can play.</p>

<h2 id="what-to-watch-carefully">What to Watch Carefully</h2>

<p>Agreeing with the broad direction, there are still a few points to hold onto so as not to be swept along.</p>

<p>First, timing. The five-year figure may slip, and the viral three-year one more so. Remembering that full self-driving has been called “ready next year” every year for over a decade, the more precisely a prophecy pins a date, the more its aim tends to be the feeling it stirs rather than the date itself. Believe the direction, but do not believe the calendar.</p>

<p>Next, the prescription. That the diagnosis of income tilting toward capital is right is a different matter from what one should buy now. The latter is usually delivered in a tone of certainty without verification. The more confident the sentence, the more one needs the habit of checking its basis separately.</p>

<p>Last, distribution. If the wealth transfer really is large, it is not a problem solved by an individual’s clever asset choices but one society has to address together. A prescription aimed only at those who can afford to buy assets risks enlarging the very problem it diagnosed. Musk himself has elsewhere floated notions like universal high income, which suggests that distribution after automation is a question at the level of institutions, not individuals.</p>

<h2 id="closing">Closing</h2>

<p>The one thing we can reliably take from this clip is this. Automation is real as a direction, and its value accrues to the ability to actually run it: for individuals, as capability that cannot be replaced; for organizations, as their own compute and data. Even if Musk’s five years is not an exact calendar, we largely agree with the direction that the relationship between intelligence and labor is changing. Only, what lies past that door is not a ticker to rush and buy but capability and infrastructure to build up slowly. That is how we reread it.</p>

<h2 id="sources">Sources</h2>

<ul>
  <li>Source video: A clip of Elon Musk in conversation with BlackRock CEO Larry Fink at the World Economic Forum in Davos. The quotes in this piece were transcribed directly from that video.</li>
  <li>Cross-checked remarks: Musk’s five-year predictions (digital intelligence exceeding the sum of human intelligence, 100 million to 1 billion humanoid robots, the economy doubling within five to seven years) were reported by multiple outlets as Davos remarks.</li>
  <li>The viral phrase “three years” is the framing of the author who quoted the clip, not Musk’s own statement.</li>
  <li>Roy Amara’s observation (short-run overestimation, long-run underestimation) is a widely cited rule of thumb in technology forecasting.</li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="culture" /><category term="AIAutomation" /><category term="FutureOfWork" /><category term="ComputeInfrastructure" /><category term="SovereignAI" /><category term="AgentEconomy" /><category term="ElonMusk" /><category term="TechPhilosophy" /><summary type="html"><![CDATA[Elon Musk's three Davos predictions got recompressed online into a warning that 'the age of the paycheck ends in three years.' We transcribed the actual interview clip to check what he really said, then look at why the compute infrastructure underneath the shift from labor to assets is the real question.]]></summary></entry><entry xml:lang="en"><title type="html">The Question Behind a 25 Billion Won Bill: Why AI Agent Costs Have Become Invisible</title><link href="https://thakicloud.github.io/en/llmops/agent-cost-observability-billing-crisis/" rel="alternate" type="text/html" title="The Question Behind a 25 Billion Won Bill: Why AI Agent Costs Have Become Invisible" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/en/llmops/agent-cost-observability-billing-crisis</id><content type="html" xml:base="https://thakicloud.github.io/en/llmops/agent-cost-observability-billing-crisis/"><![CDATA[<p>This piece is written for platform and infrastructure engineers rolling out Claude Code or AI agents across an organization, and for finance and procurement owners who have to explain next month’s AI bill. The short version: the flood of AI cost news over the past month is not really a story about “AI being expensive.” The real problem is that the bill doesn’t explain what it’s actually for. In an agent architecture, a single user request can fan out into dozens or hundreds of model calls, tool executions, and automatic retries on failure, and the final dollar amount alone gives you no way to tell which loop leaked the money. We think this observability gap is the real source of the pain the market is feeling right now.</p>

<h2 id="overview">Overview</h2>

<p>If you had to sum up the news from late June through July 2026 in one line, it would be this: using frontier models indiscriminately for every task is unsustainably expensive, and it’s hard to even trace where that cost is coming from. The clearest illustration was dramatic. A user in Korea first saw a charge attempt of roughly 2.5 billion won, then one for roughly 25 billion won. No money actually left the user’s account, since the charge exceeded the card’s limit, but the fact that an abnormal figure was pushed all the way to the card authorization stage, not just displayed as a UI glitch, is what gives the incident its weight.</p>

<p>Around the same time, news broke at other levels of the stack. An AI cost auditing firm reviewed the bills of 60 companies and alleged significant overcharging, several large enterprises began splitting frontier-model usage across cheaper models depending on the task, and reports emerged that companies in the US and Europe were migrating to Chinese open-weight models to cut costs. Interestingly, model vendors themselves began responding in a way that essentially conceded the point, that running the top-tier model on everything, all the time, isn’t sustainable. Stories from several different directions were converging on the same spot.</p>

<h2 id="what-happened-over-the-past-month">What Happened Over the Past Month</h2>

<p>The first issue to surface was the reliability of the billing system itself. According to ZDNet Korea, a charge to a Korean university student started at roughly $1.66 million and then jumped tenfold to roughly $16.62 million. Anthropic later explained that the auto-recharge amount had been set to an abnormally high figure by mistake, but the user maintained that he had never configured auto-recharge in the first place, so why that setting existed remains unexplained. The detail that the user sent more than fifteen emails to multiple departments and only received an automated reply four days later says more about the gap in the response system than about the technical error itself.</p>

<p>The second issue was the observability of agent costs. According to AI Times and The Information, cost-auditing startup Vaudit reviewed roughly $34 million in bills across 60 companies and claimed that about $1.7 million of it was overcharged. A large share of what it reviewed was Claude Code usage, and companies including Panasonic, HP, and Honda were named as clients. The patterns Vaudit pointed to included calls made on a cheap model but billed at a premium model’s rate, charges for jobs that never actually completed, and repeated automatic retries after failures, what it called a retry storm. Two caveats are worth stating clearly here. First, Anthropic pushed back, saying it does not bill for incomplete requests or error responses and that it has seen no evidence of widespread overcharging. Second, Vaudit is a commercial auditing firm that takes a cut of any refunds it wins, so these figures are best read as one party’s findings rather than an independent audit. In other words, this is currently a standoff between an auditor’s allegations and a vendor’s denial.</p>

<p>The third was how the market responded. The Information reported that companies had started splitting workloads by task, routing simple classification, summarization, and transformation jobs to cheaper models, complex coding and agentic work to frontier models, and repetitive high-volume jobs to open-weight or self-hosted models. The Financial Times reported that DoorDash, Siemens, and Airbnb had adopted DeepSeek or Moonshot-family models to cut costs. Business Insider reported that even Anthropic’s own platform executives acknowledged that shadow IT, business units adopting AI tools on their own without coordination, had driven runaway AI spending at some companies, though their prescribed fix was task-level model selection and centralized cost governance rather than usage bans or blanket budget caps. Pricing policy itself kept shifting too. Whether the latest high-performance models were included in subscriptions, and when usage-based billing would kick in, was adjusted repeatedly, and promotional deadlines kept getting extended. One example: the free tier for Claude Fable 5 was reportedly extended through July 19. For procurement teams, the bigger headache wasn’t performance, it was not being able to predict next month’s bill.</p>

<h2 id="why-agent-costs-go-unobserved">Why Agent Costs Go Unobserved</h2>

<p>There’s a single thread running through all three strands of news. In agent workloads, the distance between what the user sees and what the bill records has grown too large. A traditional API call was one request, one response, one line of cost. A coding agent or agent SDK, by contrast, expands a single instruction into planning, tool calls, file edits, verification, and automatic retries on failure. That expansion happens somewhere the user never sees, and the bill records only the sum of it, in one line.</p>

<pre><code class="language-mermaid">flowchart TB
    U["1 user request"] --&gt; P["Agent planning"]
    P --&gt; L["Execution loop"]
    L --&gt; T["Tool calls · model calls&lt;br/&gt;tens to hundreds of times"]
    T --&gt; R{"Success?"}
    R --&gt;|"Failure"| RS["Automatic retry&lt;br/&gt;(retry storm)"]
    RS --&gt; T
    R --&gt;|"Success"| ACC["Tokens · cache · tool calls&lt;br/&gt;aggregated"]
    ACC --&gt; INV["Bill: one final line item"]
    INV -.observability gap.-&gt; U
</code></pre>

<p>In this structure, most of the points where cost leaks out sit outside the user’s field of view. A retry loop quietly runs and inflates the call count, an intermediary cloud provider sits between the actual model usage and the final billing record and creates drift between the two, and one misconfigured setting, like auto-recharge, can push an abnormal amount all the way to the card authorization stage. The three stories may look unrelated, but they’re all different faces of the same observability gap. That’s why setting a monthly cap per user isn’t enough on its own. What’s actually needed is an instrumentation layer that captures per-model cost, per-session tokens, cache tokens, tool call counts, failure and retry costs, and day-over-day anomaly rates centrally, at the moment each call happens. Without observability there’s no control, and without control the bill will always be the document that surprises you after the fact.</p>

<h2 id="implications-for-thakiclouds-products">Implications for ThakiCloud’s Products</h2>

<p>This is a problem that two of ThakiCloud’s products each target from a different angle. Because the infrastructure lens and the agent lens complement each other here, we apply both to this topic.</p>

<p><strong>The ai-platform lens: owning repetitive workloads is the answer.</strong> The market has arrived at a clear conclusion. Running even trivial tasks on frontier models is unsustainably expensive, and self-hosting an open-weight model is the economical choice for repetitive, high-volume work. ThakiCloud’s ai-platform is Kubernetes-based AI/ML infrastructure built for exactly that gap. It queues GPUs with Kueue to push utilization higher, serves open-weight models through vLLM, and separates usage and billing by department through multi-tenant isolation. Where usage-based API pricing produces unpredictable bills, self-hosting builds a structure on top of a fixed GPU cost where the unit price doesn’t spike even as usage grows. Unlike external usage-based pricing that keeps shifting, on-premises and sovereign deployments turn cost predictability itself into an asset. And the fact that data never leaves the organization is an additional value for organizations facing heavy domestic regulatory and security requirements.</p>

<p><strong>The Paxis lens: making every agent action auditable.</strong> The core of the observability gap was the agent loop, and that is precisely the territory Paxis addresses. Paxis is ThakiCloud’s Agent-Native Cloud control plane, running on top of ai-platform, that treats Skills, Tools, Policies, and Audit Logs as first-class resources. Which skill an agent invoked, through which tool, how many times, and in which sandbox it ran, all of it is captured in an audit log. In this structure, a retry storm can’t quietly inflate a bill: the retry loop shows up directly in the audit trail, and policy gates block calls that cross a threshold. A design that selects from more than 960 skills using BM25, runs them in isolated sandboxes, and routes every action through policy and audit is a structural answer to exactly the problem of not being able to tell, from the bill alone, which loop generated the cost. Low-cost serving through ai-platform makes agents economical, and action-level observability through Paxis makes that economics predictable. That’s how the two lenses fit together.</p>

<h2 id="limitations-and-counterarguments">Limitations and Counterarguments</h2>

<p>In the interest of balance, let’s state the counterargument clearly. First, frontier models aren’t automatically wasteful. According to the Wall Street Journal, companies like Shopify believe that for complex coding and multi-step agent work, a frontier model’s higher price can be justified if it saves enough engineering time. Spotify and Twilio, by contrast, are weighing more carefully whether a marginal performance gain justifies the added cost. The takeaway isn’t “abandon frontier models,” it’s “split the workload by task difficulty.” Self-hosting isn’t a universal answer either. Pushing tasks that require the highest level of reasoning down to open-weight models degrades quality, and it introduces a new operational burden of GPU operations, model updates, and security patching.</p>

<p>Second, the overbilling figures cited in this piece are not settled facts. Vaudit’s claims come from a commercial auditing firm, and Anthropic has denied them, so the accurate reading right now is that the two sides’ positions are in direct conflict. In the 25-billion-won billing incident too, no money actually changed hands, and no technical explanation for why the auto-recharge setting existed has been made public. The conclusion we draw from this news isn’t aimed at any particular vendor. It’s a principle: in the agent era, whichever vendor you use, cost observability and governance need to be secured on the user’s side. Choosing a good model and governing that model are two separate problems, and the news of the past month simply exposed the fact that the latter has been an empty seat all along.</p>

<h2 id="sources">Sources</h2>

<ul>
  <li><a href="https://zdnet.co.kr/view/?no=20260709165452">ZDNet Korea, “Korean User Hit With 25 Billion Won Payment Request, Anthropic Billing Error Controversy” (2026-07-09)</a></li>
  <li><a href="https://zdnet.co.kr/view/?no=20260716093004">ZDNet Korea, “Anthropic’s 25 Billion Won Charge Turns Out to Be an Auto-Recharge Configuration Error” (2026-07-16)</a></li>
  <li><a href="https://www.aitimes.com/news/articleView.html?idxno=212155">AI Times, “Anthropic Faces ‘AI Overbilling’ Controversy, Charged for Failed Jobs Too” (2026-06)</a></li>
  <li><a href="https://www.theinformation.com/titv/fedld">The Information, report on enterprises adopting AI cost controls and model diversification (2026-06-23)</a></li>
  <li><a href="https://www.ft.com/content/9c8ff45b-7c20-4c2e-93c9-c52339ffdcee">Financial Times, “Companies turn to Chinese AI models to cut costs” (2026-07)</a></li>
  <li><a href="https://www.businessinsider.com/anthropic-ai-costs-responses-routers-2026-7">Business Insider, “Anthropic Official Warns Against ‘Wrong’ AI Cost Response” (2026-07-15)</a></li>
  <li><a href="https://www.wsj.com/cio-journal/meet-the-companies-shelling-out-for-top-ai-models-e1fe3375">The Wall Street Journal, “Meet the Companies Shelling Out for Top AI Models” (2026-07)</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="llmops" /><category term="LLMOps" /><category term="FinOps" /><category term="AgentCost" /><category term="CostObservability" /><category term="ModelRouting" /><category term="self-hosting" /><category term="Paxis" /><category term="AIInfrastructure" /><summary type="html"><![CDATA[From a 25 billion won bill sent to a user in Korea to allegations of overbilling at 60 companies, the past month of AI cost news points to a single blind spot. In the age of agents, where a single request can trigger hundreds of model calls, the bill no longer explains what you actually paid for. Here is how we think that observability gap should be closed.]]></summary></entry><entry xml:lang="en"><title type="html">A 122B Model on a 24GB Card? We Dissected ATSInfer, Which Slices llama.cpp at Tensor Granularity</title><link href="https://thakicloud.github.io/en/llmops/atsinfer-hybrid-cpu-gpu-tensor-scheduling/" rel="alternate" type="text/html" title="A 122B Model on a 24GB Card? We Dissected ATSInfer, Which Slices llama.cpp at Tensor Granularity" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/en/llmops/atsinfer-hybrid-cpu-gpu-tensor-scheduling</id><content type="html" xml:base="https://thakicloud.github.io/en/llmops/atsinfer-hybrid-cpu-gpu-tensor-scheduling/"><![CDATA[<p>This post is for engineers weighing whether to self-serve a large model on a single consumer GPU, and for infra owners deciding how much to trust the “run 120B on 24GB” tweets going around. Up front: the core idea of ATSInfer (arXiv:2607.10183), released by researchers at Nanjing University, is simple and persuasive. Where prior offloading moved things in chunks at the granularity of a “layer” or an “expert,” ATSInfer slices down to <strong>individual tensors</strong>. That said, the headline “up to 3.29x” rests on a few premises, and the code is not yet public. We did not reproduce an RTX 4090 running a 120B-class model here, so every number in this post is a <strong>value reported by the paper</strong>, stated as such.</p>

<h2 id="overview">Overview</h2>

<p>Anyone who has run a local LLM hits the same wall: if the model weights are larger than GPU memory, the overflow spills to CPU memory. llama.cpp’s <code class="language-plaintext highlighter-rouge">-ngl</code> flag (how many layers to place on the GPU) does exactly this. The problem is that it cuts only at <strong>layer granularity</strong>. A single layer mixes tensors of very different character (attention weights, FFN weights, normalization params), yet they are handled all-or-nothing: either the whole thing goes to the GPU or it stays on the CPU.</p>

<p>Why this chunked placement loses is simple. The same 1GB in VRAM might make one tensor 10x faster and another only 2x faster. VRAM is a scarce resource, and chunked cuts prevent you from picking the tensors with the highest gain per GB. ATSInfer targets exactly this. It profiles each tensor’s CPU and GPU performance and fills VRAM starting from the tensors with the <strong>highest speed gain per GB</strong>. If the recent ktransformers was an “expert-level” trick that pushes MoE experts to the CPU (see our <a href="/en/llmops/ktransformers-moe-offload-28x-validation/">related post: reproducing the ktransformers 28x</a>), ATSInfer is the finer, “tensor-level” generalization. Notably, it applies to dense models, not just MoE.</p>

<h2 id="what-is-this-technology">What is this technology</h2>

<p>ATSInfer is a hybrid CPU-GPU inference system built as an extension to llama.cpp in roughly 15,000 lines of C++. As the name says, “Automated Tensor Scheduling” is the core, and three mechanisms interlock.</p>

<pre><code class="language-mermaid">flowchart TB
    A["Model weights&lt;br/&gt;(RAM, exceed VRAM capacity)"] --&gt; B["Per-tensor profiling&lt;br/&gt;measure speed gain per GB"]
    B --&gt; C{"Static placement&lt;br/&gt;highest-gain tensors to VRAM first"}
    C --&gt;|"high-gain tensors"| D["Resident in VRAM"]
    C --&gt;|"low-gain tensors"| E["Resident in RAM"]
    D --&gt; F["Load-aware dynamic transfer&lt;br/&gt;promote/demote by runtime load"]
    E --&gt; F
    F --&gt; G["Asynchronous CPU-GPU coordination&lt;br/&gt;overlap compute and PCIe transfer"]
    G --&gt; H["Token output&lt;br/&gt;prefill · decode"]
</code></pre>

<p><strong>First, static tensor placement.</strong> Before loading the model, it benchmarks how much each tensor speeds up on the GPU, then places tensors on the GPU in order of “most speed returned per GB of VRAM used.” This is close to a knapsack optimization and directly exploits the tensor-level heterogeneity that chunked placement ignored.</p>

<p><strong>Second, load-aware dynamic transfer.</strong> Static placement alone is not enough. During real inference, load shifts moment to moment with batch size, context length, and concurrency. ATSInfer promotes a given tensor from RAM to GPU, or demotes it, based on the runtime situation. If static placement is the starting line, dynamic transfer is changing lanes while driving.</p>

<p><strong>Third, asynchronous CPU-GPU coordination.</strong> It overlaps CPU compute, GPU compute, and the PCIe transfer that connects them. A naive implementation leaves the GPU idling while it waits on CPU work or data movement; this coordination layer fills that idle time. The paper reports this raises average GPU SM (streaming multiprocessor) utilization by about 70%.</p>

<h2 id="results-the-paper-reports">Results the paper reports</h2>

<p>Again: the numbers below are <strong>values reported by the paper</strong>, not something we reproduced. ATSInfer’s code is not yet public (even the tweet chatter said “researchers, please share the code with the llama.cpp team”), and a 120B-class model plus an RTX 4090 is hard to reproduce in the sandbox behind this post. So instead of reproduction, we focus on <strong>structural analysis and implications</strong>.</p>

<p>The paper’s headline: versus existing hybrid systems (including llama.cpp’s layer-level offloading), prefill (throughput to first token) improves by up to 1.94x, and decode (tokens generated per second) by up to 3.29x.</p>

<p><img src="/assets/images/atsinfer-hybrid-cpu-gpu-tensor-scheduling-results.png" alt="Max speedup ATSInfer reports in the paper" /></p>

<p>The setup is an RTX 4090 (24GB) and RTX 3060 system with 64GB RAM, and the validated models are:</p>

<ul>
  <li>Llama 3.1-70B (INT4)</li>
  <li>Qwen3-Next-80B-A3B (INT4)</li>
  <li>Qwen3.5-122B-A10B (INT4)</li>
  <li>GPT-OSS-120B (MXFP4)</li>
</ul>

<p>So the central claim is running a 122B-parameter model (an MoE with far fewer active parameters) on a single 24GB card. Read two things separately here. First, “3.29x” is a <strong>maximum</strong> under specific conditions, not an average across every model and batch. Second, the gain fundamentally comes not from “fitting what didn’t fit into the GPU” but from “making the unavoidable CPU-GPU traffic smarter.” The physics that PCIe bandwidth is the bottleneck stays the same, so ATSInfer’s contribution is using that bandwidth without waste and reducing GPU idle time.</p>

<h2 id="implications-for-thakicloud-products">Implications for ThakiCloud products</h2>

<p>ThakiCloud’s <strong>ai-platform</strong> is an AI/ML infrastructure that serves models across diverse customer environments on Kubernetes and Kueue. Tensor-level scheduling like ATSInfer aligns with a trend we watch closely.</p>

<p>First, <strong>the economics of on-premises and sovereign environments.</strong> In settings where data cannot leave the premises, such as domestic public-sector and financial customers, models must run on owned GPUs. If a few consumer GPUs can carry a mid-to-large model instead of a rack of eight H100s, initial CAPEX drops dramatically. What ATSInfer’s experiments show is that the premise “if VRAM is short you must simply buy more GPUs” can be substantially relaxed through tensor placement optimization. The cost, of course, is reduced throughput, so it is unsuitable for latency-critical workloads. Judging that trade-off per workload is the job of our serving layer.</p>

<p>Second, <strong>coupling with multi-tenant scheduling.</strong> ATSInfer’s “load-aware dynamic transfer” is tensor movement within a single node, but the idea holds at cluster scale too. When Kueue queues and allocates GPU resources, a policy that decides which request to handle at which precision and offloading profile based on load is an area we already think about. Just as tensor-level profiling squeezes resource gains within a node, the cluster scheduler does the same across nodes.</p>

<p>Third, <strong>redefining the cost-quality curve.</strong> In our <a href="/en/llmops/ktransformers-moe-offload-28x-validation/">ktransformers reproduction post</a> we showed by direct measurement that headline figures like “28x” rest on hidden premises. ATSInfer’s “3.29x” deserves the same lens. Not a marketing number, but the value that emerges on our customers’ actual models, actual batches, and actual SLAs, is what we verify. Competitiveness at low serving cost ultimately comes from accumulating exactly this kind of verification.</p>

<h2 id="limits-and-counterarguments">Limits and counterarguments</h2>

<p>The biggest limit is <strong>the code being unreleased.</strong> Whether the paper’s numbers reproduce, and whether they hold on other hardware and models, can only be checked once code appears. A 15,000-line C++ extension is also a nontrivial maintenance and upstream-merge task. A fork that never merges falls behind upstream changes over time and loses value.</p>

<p>Second, <strong>the condition-dependence of the gains.</strong> The effect of tensor-level placement depends heavily on CPU performance, RAM bandwidth, and PCIe generation. The paper’s experiments assume 64GB RAM; if RAM is short, there is no room to keep tensors on the CPU at all. On PCIe 3.0 systems, transfer likely becomes the bottleneck and shrinks the gain substantially ([estimate], the paper does not spell out a generation-by-generation comparison).</p>

<p>Third, <strong>the inherent ceiling on decode optimization.</strong> Decode is a memory-bound task. However well you schedule, weights outside VRAM must be accessed somehow every token, so it is inevitably slower than pure VRAM residency. What ATSInfer does is “minimize how much slower,” not “eliminate the slowdown.” Building the opposing case: for production serving that truly needs low latency and high throughput, using a GPU that holds the whole model in VRAM is still the right call. ATSInfer shines in the development, evaluation, and small-batch regime where you cannot afford that GPU, or do not need it.</p>

<p>Even so, the direction has clear value. Squeezing resource utilization through software rather than adding hardware fits a platform like ours particularly well, one that treats on-prem, cost-efficiency, and self-hosting as its weapons. Once the code is public, we plan to fold it into our serving benchmarks and measure the real numbers ourselves.</p>

<h2 id="sources">Sources</h2>

<ul>
  <li>ATSInfer paper: <a href="https://arxiv.org/abs/2607.10183">arXiv:2607.10183, Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices</a></li>
  <li>Related post: <a href="/en/llmops/ktransformers-moe-offload-28x-validation/">A $400K rack on 24GB? We reproduced the ktransformers 28x</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="llmops" /><category term="ATSInfer" /><category term="llama.cpp" /><category term="CPU-offloading" /><category term="GPU" /><category term="LLM-serving" /><category term="LLMOps" /><category term="quantization" /><category term="infrastructure" /><summary type="html"><![CDATA[Instead of offloading whole layers or experts, ATSInfer places individual tensors across CPU and GPU. The paper reports running models far beyond VRAM on a single RTX 4090 while pushing decode up to 3.29x. We read arXiv:2607.10183 and unpacked how it works and what it means for serving.]]></summary></entry><entry xml:lang="ko"><title type="html">Kimi Code CLI 뜯어보기: 오픈소스 터미널 에이전트가 ACP로 에디터를 삼키는 법</title><link href="https://thakicloud.github.io/ko/agentops/kimi-code-cli-acp-open-source-agent/" rel="alternate" type="text/html" title="Kimi Code CLI 뜯어보기: 오픈소스 터미널 에이전트가 ACP로 에디터를 삼키는 법" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/ko/agentops/kimi-code-cli-acp-open-source-agent</id><content type="html" xml:base="https://thakicloud.github.io/ko/agentops/kimi-code-cli-acp-open-source-agent/"><![CDATA[<p>지난주 문샷AI가 오픈웨이트 모델 Kimi K3를 공개하면서 코딩 리더보드 1위를 가져갔습니다. 그런데 모델보다 개발자 워크플로에 더 직접 닿는 물건이 조용히 같이 나왔습니다. <strong>Kimi Code CLI</strong>, 문샷이 MIT 라이선스로 공개한 오픈소스 터미널 코딩 에이전트입니다. 링크드인 타임라인에는 “클로드 코드에 없는 기능을 준다”는 소개가 돌았습니다. 저희는 그 문장을 그대로 옮기지 않고 공식 저장소와 문서를 직접 확인했습니다. 결론부터 말하면 소개의 절반은 사실이고 절반은 과장입니다. 그리고 진짜 흥미로운 지점은 홍보 문구가 강조하지 않은 곳에 있었습니다.</p>

<p>이 글은 Kimi Code CLI가 무엇인지, 무엇을 실제로 제공하는지, 그리고 K8s 기반 AI 플랫폼을 운영하는 저희 관점에서 왜 눈여겨볼 만한지를 정리합니다. 특히 Agent Client Protocol이라는 개방 표준이 왜 에이전트 생태계의 판을 바꾸는 조각인지에 지면을 많이 할애했습니다.</p>

<h2 id="kimi-code-cli는-무엇인가">Kimi Code CLI는 무엇인가</h2>

<p>Kimi Code CLI는 터미널에서 도는 에이전트형 코딩 도구입니다. 클로드 코드나 제미나이 CLI, 코덱스 CLI와 같은 계열입니다. 공식 저장소는 <a href="https://github.com/MoonshotAI/kimi-code">MoonshotAI/kimi-code</a>이며, 이전 프로젝트인 <a href="https://github.com/MoonshotAI/kimi-cli">MoonshotAI/kimi-cli</a>가 여기로 진화하면서 기존 세션과 설정이 이어집니다. 두 저장소 모두 문샷 공식이며, 이름이 비슷한 서드파티 프로젝트와 혼동하지 않도록 주의가 필요합니다.</p>

<p>여기서 첫 번째 정정이 필요합니다. 이 CLI의 정식 이름은 “Kimi K3용 CLI”가 아니라 <strong>Kimi Code CLI</strong>입니다. 모델에 종속되지 않는 도구이고, 기본값으로 문샷의 코딩 특화 모델인 Kimi K2.7 Code를 붙여 쓰지만 설정으로 K3를 포함해 다른 모델로도 전환할 수 있습니다. K3는 CLI가 붙일 수 있는 여러 모델 중 하나이지, CLI가 K3 전용으로 만들어진 것은 아닙니다. K3 자체는 문샷이 2026년 7월 16일 공개한 2조8000억 매개변수급 오픈 MoE 모델로, Kimi Delta Attention과 최대 100만 토큰 컨텍스트를 내세웁니다. 이 부분은 CNBC, 블룸버그, 포브스 등 주요 매체가 함께 보도했습니다.</p>

<p>전체 그림을 먼저 세워두면 이후 세부가 훨씬 잘 붙습니다. Kimi Code CLI가 한쪽으로는 MCP 클라이언트로 도구와 데이터에 연결되고, 다른 한쪽으로는 ACP 서버로 에디터에 연결되는 이중 역할이 핵심입니다.</p>

<pre><code class="language-mermaid">flowchart TB
    subgraph EDITOR["개발자 에디터 (ACP 클라이언트)"]
        ZED["Zed"]
        JB["JetBrains 계열"]
        VSC["VS Code / Neovim"]
    end
    ACP["Agent Client Protocol&lt;br/&gt;JSON-RPC over stdio"]
    subgraph CLI["Kimi Code CLI (에이전트 코어)"]
        MAIN["메인 에이전트&lt;br/&gt;대화 히스토리 유지"]
        SUB["서브에이전트&lt;br/&gt;coder · explore · plan&lt;br/&gt;각자 격리 컨텍스트"]
    end
    MODEL["모델 계층&lt;br/&gt;Kimi K2.7 Code / K3&lt;br/&gt;또는 OpenAI 호환 엔드포인트"]
    subgraph MCP["MCP 서버 (도구 · 데이터)"]
        T1["Context7"]
        T2["Chrome DevTools"]
        T3["사내 커넥터"]
    end

    EDITOR --&gt; ACP
    ACP --&gt;|kimi acp| MAIN
    MAIN --&gt; SUB
    MAIN --&gt;|추론 요청| MODEL
    SUB --&gt;|추론 요청| MODEL
    MAIN --&gt;|도구 호출| MCP
</code></pre>

<h2 id="서브에이전트-컨텍스트를-나눠서-메인을-깨끗하게">서브에이전트: 컨텍스트를 나눠서 메인을 깨끗하게</h2>

<p>Kimi Code CLI는 세 종류의 내장 서브에이전트를 제공합니다. <strong>coder</strong>는 파일을 읽고 쓰고 명령을 실행해 실제 변경을 반영하는 범용 엔지니어링 담당입니다. <strong>explore</strong>는 읽기 전용으로 코드베이스를 훑는 탐색 담당입니다. <strong>plan</strong>은 셸 명령 없이 구현 계획과 아키텍처 설계만 내놓는 담당입니다. 이 구분은 공식 문서 <a href="https://moonshotai.github.io/kimi-code/en/customization/agents.html">Agents and Sub-Agents</a>에 명시되어 있습니다.</p>

<p>핵심은 이름이 아니라 컨텍스트 격리입니다. 각 서브에이전트는 완전히 독립된 컨텍스트 윈도우를 가지며, 메인 에이전트가 명시적으로 넘긴 작업 설명만 봅니다. 메인의 대화 히스토리는 서브에이전트에 노출되지 않고, 서브에이전트가 돌리는 중간 추론과 도구 호출 로그도 메인 히스토리에 섞이지 않습니다. 서브에이전트는 최종 결론만 반환합니다. 긴 세션에서 메인 컨텍스트가 로그로 부풀지 않고 얇게 유지되는 이유가 여기 있습니다. 백그라운드 실행과 병렬 실행도 지원해서, 여러 탐색 작업을 동시에 돌리고 완료되면 자동으로 결과가 돌아옵니다.</p>

<p>이 패턴은 저희에게 낯설지 않습니다. 이 블로그를 운영하는 내부 오케스트레이션 하네스도 탐색은 저비용 서브에이전트에 위임하고 요약만 회수해 메인 컨텍스트를 보호합니다. 컨텍스트 위생이 곧 비용이자 품질이라는 원칙은 도구가 달라도 같습니다.</p>

<h2 id="mcp-json을-손으로-고치지-않는-설정-경험">MCP: JSON을 손으로 고치지 않는 설정 경험</h2>

<p>Model Context Protocol 연동은 두 경로로 관리합니다. 첫째, CLI 서브커맨드입니다. <code class="language-plaintext highlighter-rouge">kimi mcp add</code>, <code class="language-plaintext highlighter-rouge">kimi mcp list</code>, <code class="language-plaintext highlighter-rouge">kimi mcp remove</code>, <code class="language-plaintext highlighter-rouge">kimi mcp authorize</code>로 서버를 다룹니다. 예를 들어 HTTP 트랜스포트로 문서 검색 서버를 붙이거나, stdio 트랜스포트로 브라우저 자동화 서버를 붙일 수 있습니다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># HTTP 트랜스포트 (OAuth 옵션 지원)</span>
kimi mcp add <span class="nt">--transport</span> http context7 https://mcp.context7.com/mcp

<span class="c"># stdio 트랜스포트로 로컬 프로세스 연결</span>
kimi mcp add <span class="nt">--transport</span> stdio chrome-devtools <span class="nt">--</span> npx chrome-devtools-mcp@latest
</code></pre></div></div>

<p>둘째, TUI 안에서 쓰는 대화형 슬래시 명령 <code class="language-plaintext highlighter-rouge">/mcp-config</code>입니다. JSON 설정 파일을 직접 편집하지 않고 서버를 추가, 수정, 인증할 수 있습니다. <code class="language-plaintext highlighter-rouge">/mcp</code>는 현재 연결된 서버와 로드된 도구 목록을 보여줍니다. 링크드인 소개가 강조한 “JSON을 직접 수정할 필요가 없다”는 부분은 사실입니다. 다만 이 편의성 자체가 클로드 코드에 없는 것은 아닙니다. 이 지점은 뒤에서 다시 정리합니다. 관련 문서는 <a href="https://moonshotai.github.io/kimi-cli/en/customization/mcp.html">MCP 설정</a>에 있습니다.</p>

<h2 id="agent-client-protocol-이-도구에서-가장-중요한-조각">Agent Client Protocol: 이 도구에서 가장 중요한 조각</h2>

<p>여기가 이 글에서 가장 흥미로운 부분입니다. Agent Client Protocol, 줄여서 ACP는 Zed 에디터 팀이 만든 개방 표준입니다. Apache 라이선스이고, JSON-RPC 2.0을 stdio 위에서 주고받습니다. 에디터가 에이전트를 자식 프로세스로 띄우고 표준 입출력으로 통신하는 방식으로, 전송 메커니즘 자체는 언어 서버 프로토콜과 동일합니다.</p>

<p>비유가 이해를 크게 돕습니다. LSP가 등장하기 전에는 에디터마다 언어마다 별도 통합을 만들어야 했습니다. LSP는 이 M 곱하기 N 문제를 M 더하기 N으로 바꿨습니다. 에디터 하나가 표준만 구현하면 누가 만든 언어 서버든 그 덕을 봅니다. ACP는 정확히 같은 일을 에이전트에 합니다. 에디터 하나가 ACP를 구현하면 누가 만든 에이전트든 그 에디터에 표준 방식으로 꽂힙니다. 이 개념은 <a href="https://zed.dev/acp">Zed의 ACP 소개</a>와 <a href="https://blog.marcnuri.com/agent-client-protocol-acp-introduction">마크 누리의 해설 글</a>에서 확인할 수 있습니다.</p>

<p>MCP와 헷갈리기 쉬운데 방향이 반대입니다. MCP는 에이전트에서 도구와 데이터로 향합니다. 이때 에이전트가 MCP 클라이언트입니다. ACP는 에디터에서 에이전트로 향합니다. 이때 에이전트가 ACP 서버이고 에디터가 ACP 클라이언트입니다. 같은 에이전트가 한쪽으로는 MCP 클라이언트, 다른 쪽으로는 ACP 서버 역할을 동시에 맡습니다. 앞의 다이어그램이 이 이중 역할을 그린 이유입니다.</p>

<p>Kimi Code CLI는 <code class="language-plaintext highlighter-rouge">kimi acp</code> 서브커맨드로 이 프로토콜을 별도 설치 없이 네이티브로 지원합니다. Zed는 네이티브로, JetBrains는 플러그인으로 연결되고, Zed의 ACP 레지스트리 기준으로는 여러 에디터 통합이 이미 올라와 있습니다. 개발자는 익숙한 에디터를 떠나지 않고 그 안에서 Kimi 세션을 몰 수 있습니다.</p>

<h2 id="이미지와-비디오-입력-정확히-어디까지-되나">이미지와 비디오 입력, 정확히 어디까지 되나</h2>

<p>링크드인 소개는 “화면 캡처를 그대로 입력으로 전달할 수 있다”고 적었습니다. 여기에 정정이 필요합니다. 문샷이 정면에 내세우는 기능은 정적 스크린샷이 아니라 <strong>화면 녹화 영상 입력</strong>입니다. 저장소 설명은 화면 녹화나 데모 클립을 채팅에 떨어뜨리면 에이전트가 말로 설명하기 어려운 동작을 직접 보고 이해한다고 표현합니다. 물론 CLI 입력창에서 이미지 붙여넣기도 지원하며, 기본 모델 Kimi K2.7 Code가 4억 매개변수 비전 인코더 MoonViT를 갖춘 네이티브 멀티모달이라 텍스트, 이미지, 비디오를 모두 받습니다. 다만 커스텀 모델을 붙일 때는 해당 모델의 modalities에 이미지 지원을 명시해야 정상 동작합니다. 요약하면 이미지 입력이 되긴 하지만, 진짜 차별 포인트로 홍보되는 것은 영상 입력이며 “스크린샷”이라는 표현은 다소 부정확합니다.</p>

<h2 id="설치는-실제로-세-단계">설치는 실제로 세 단계</h2>

<p>설치 흐름은 소개대로 간결합니다. 아래 명령은 <a href="https://moonshotai.github.io/kimi-cli/en/guides/getting-started.html">공식 시작 가이드</a> 기준이며, 저희 사내 샌드박스가 해당 배포 도메인에 접근 권한이 없어 직접 실행 로그는 남기지 않았습니다. 따라서 벤치마크 수치는 만들지 않고, 검증된 명령만 옮깁니다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># 1) 설치 스크립트 실행 (uv를 함께 설치)</span>
curl <span class="nt">-LsSf</span> https://code.kimi.com/install.sh | bash

<span class="c"># 2) 프로젝트 디렉터리에서 실행</span>
kimi

<span class="c"># 3) 인증 설정</span>
/login
</code></pre></div></div>

<p>macOS는 <code class="language-plaintext highlighter-rouge">brew install kimi-code</code>, 윈도우는 파워셸 스크립트도 제공합니다. 소스에서 개발하려면 Node 24.15 이상과 pnpm이 필요합니다. 라이선스는 MIT라서, 코드를 열어 보고 포크하고 사내 배포하는 데 제약이 적습니다.</p>

<h2 id="모델과-프로바이더-개방성">모델과 프로바이더 개방성</h2>

<p>컨텍스트 길이는 K2.6 계열이 최대 25만6000 토큰이고, K3는 문샷 마케팅 기준 최대 100만 토큰입니다. 더 중요한 것은 프로바이더 개방성입니다. <code class="language-plaintext highlighter-rouge">~/.kimi-code/config.toml</code>에서 OpenAI 호환 엔드포인트, Anthropic API 키, 구글 GenAI나 Vertex AI까지 다중 프로바이더로 등록할 수 있습니다. CLI가 특정 모델에 락인되지 않는다는 뜻입니다. 서드파티 추론 모델의 reasoning_content 필드도 자동 처리합니다. 관련 문서는 <a href="https://moonshotai.github.io/kimi-cli/en/configuration/providers.html">Providers and models</a>에 있습니다.</p>

<h2 id="클로드-코드에는-없는-기능인가-정직한-비교">클로드 코드에는 없는 기능인가: 정직한 비교</h2>

<p>가장 널리 퍼진 소개 문구는 “클로드 코드에는 없는 기능을 제공한다”였습니다. 검증 결과 이 프레이밍은 대부분 과장입니다.</p>

<p>서브에이전트와 격리 컨텍스트는 클로드 코드도 서브에이전트 기능으로 같은 방식을 제공합니다. MCP도 클로드 코드가 stdio, SSE, HTTP 전송을 이미 성숙하게 지원합니다. 이미지 붙여넣기도 클로드 코드에 있습니다. 이 세 가지는 차별점이 아닙니다.</p>

<p>진짜 차이는 두 곳에 있습니다. 첫째, ACP 지원 방식입니다. Kimi Code CLI는 <code class="language-plaintext highlighter-rouge">kimi acp</code> 서브커맨드로 CLI 자체에 ACP를 1차 기능으로 내장합니다. 반면 클로드 코드는 Zed가 만든 별도 어댑터 패키지를 거쳐 베타로 연결됩니다. 사용자 입장에서 전자는 도구를 켜면 바로 되고, 후자는 브릿지를 하나 더 얹어야 합니다. 둘째, 모델 개방성입니다. Kimi는 오픈웨이트 K 시리즈에 멀티 프로바이더 전환까지 열려 있는 반면, 클로드 코드는 앤트로픽 모델 전용입니다. 여기서 셀프호스팅 가능성이라는 세 번째 차이가 파생됩니다. Kimi는 오픈소스 CLI에 오픈웨이트 모델이라 온프렘 서빙이 가능하지만, 클로드 코드는 CLI는 열려 있어도 모델이 API 전용입니다. 관련 근거는 <a href="https://zed.dev/blog/claude-code-via-acp">Zed의 클로드 코드 ACP 베타 글</a>에서 확인할 수 있습니다.</p>

<h2 id="thakicloud-제품-적용-시사점">ThakiCloud 제품 적용 시사점</h2>

<p>이 주제는 에이전트 도구인 동시에 개방 모델과 온프렘이라는 인프라 축을 건드립니다. 그래서 두 렌즈를 함께 씁니다.</p>

<p>Paxis 렌즈로 보면 Kimi Code CLI의 구조가 저희 제품의 설계 방향과 상당히 겹칩니다. Paxis는 ThakiCloud의 Agent-Native Cloud 제어 평면으로, 스킬과 도구와 정책과 감사 로그를 일급 리소스로 다룹니다. Kimi의 coder, explore, plan 서브에이전트가 격리 컨텍스트에서 병렬로 도는 방식은 Paxis의 스킬 하네스가 960개 넘는 스킬을 BM25로 선택해 격리 샌드박스에서 실행하는 방식과 같은 철학을 공유합니다. 특히 ACP는 벤더 중립 표준이라는 점에서 Paxis에 직접적인 기회입니다. 저희가 배포하는 어떤 에이전트든, 자체 파인튜닝 모델을 얹은 것까지 포함해서, ACP를 구현하면 고객의 Zed나 JetBrains 같은 개발 에디터에 표준 방식으로 꽂힐 수 있습니다. MCP 커넥터로 데이터에 연결하고 ACP로 에디터에 연결하는 이중 표준 조합은 정확히 저희가 지향하는 통합 그림입니다.</p>

<p>ai-platform 렌즈로 보면 개방성이 곧 배포 자유입니다. 오픈웨이트 K 시리즈를 저희 클러스터의 Kueue GPU 스케줄링과 vLLM 서빙 위에 올리고, CLI를 사내 엔드포인트로 라우팅하면 외부 API 의존이나 데이터 반출 없이 사내 코딩 에이전트를 구축할 수 있습니다. 코드가 외부로 나가면 안 되는 금융이나 공공 영역, 그리고 국정원 요구사항 같은 온프렘 보안 요건과 정합합니다. 능력이 흔해지고 싸질수록 기업이 실제로 지불하는 것은 통제된 실행 환경이라는 관점은 저희가 이전 글에서도 다룬 주제입니다. Kimi Code CLI는 그 실행 계층을 오픈소스로 열어 놓았다는 점에서 의미가 있습니다.</p>

<h2 id="한계-및-반론">한계 및 반론</h2>

<p>몇 가지는 냉정하게 봐야 합니다. 첫째, 서드파티 딥다이브에서 언급되는 내부 엔진 명칭이나 계층 구조는 공식 문서에 없는 표현이라 리버스 엔지니어링일 가능성이 있습니다. 사실로 인용하기 전에 공식 문서를 기준으로 삼는 편이 안전합니다. 둘째, ACP 경로가 다른 연결 방식보다 응답 품질이 낫다는 커뮤니티 보고가 있지만 벤치마크가 아니라 체감입니다. 검증된 수치가 아닙니다. 셋째, 오픈웨이트라 해도 2조8000억 매개변수급 모델을 온프렘에서 실제로 서빙하려면 상당한 GPU 자원이 필요합니다. 개방성이 곧 손쉬운 셀프호스팅을 뜻하지는 않습니다. 작은 팀에는 API 경로가 여전히 현실적입니다. 넷째, 도구 생태계의 성숙도와 안정성은 클로드 코드나 코덱스 CLI가 앞서 있을 수 있습니다. 오픈소스라는 사실이 곧 프로덕션 준비 완료를 의미하지는 않습니다.</p>

<p>그럼에도 개방 표준 위에서 에이전트와 에디터가 느슨하게 결합되는 방향은 분명한 흐름입니다. 특정 벤더의 CLI에 종속되지 않고, 모델과 에디터를 각각 갈아 끼울 수 있는 세계가 개발자에게 더 유리합니다. Kimi Code CLI는 그 세계를 앞당기는 조각 중 하나입니다.</p>

<h2 id="출처">출처</h2>

<ul>
  <li><a href="https://github.com/MoonshotAI/kimi-code">MoonshotAI/kimi-code (공식 저장소)</a></li>
  <li><a href="https://github.com/MoonshotAI/kimi-cli">MoonshotAI/kimi-cli (이전 저장소)</a></li>
  <li><a href="https://moonshotai.github.io/kimi-cli/en/guides/getting-started.html">Kimi Code CLI 시작 가이드</a></li>
  <li><a href="https://moonshotai.github.io/kimi-code/en/customization/agents.html">Agents and Sub-Agents 문서</a></li>
  <li><a href="https://moonshotai.github.io/kimi-cli/en/customization/mcp.html">MCP 설정 문서</a></li>
  <li><a href="https://zed.dev/acp">Zed - Agent Client Protocol</a></li>
  <li><a href="https://blog.marcnuri.com/agent-client-protocol-acp-introduction">ACP: The LSP for AI Coding Agents</a></li>
  <li><a href="https://zed.dev/blog/claude-code-via-acp">Zed - Claude Code via ACP (베타)</a></li>
  <li><a href="https://www.marktechpost.com/2026/07/16/moonshot-ai-releases-kimi-k3-a-2-8-trillion-parameter-open-moe-model-with-kimi-delta-attention-and-1m-context/">MarkTechPost - Kimi K3 공개</a></li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="agentops" /><category term="agentops" /><category term="kimi" /><category term="moonshot" /><category term="coding-agent" /><category term="mcp" /><category term="agent-client-protocol" /><category term="paxis" /><category term="thakicloud" /><summary type="html"><![CDATA[문샷AI가 Kimi K3와 함께 공개한 오픈소스 코딩 CLI를 실제 문서와 저장소 기준으로 분석합니다. coder/explore/plan 서브에이전트, 대화형 MCP 설정, 그리고 진짜 차별점인 Agent Client Protocol 네이티브 지원까지, '클로드 코드에 없는 기능'이라는 홍보 문구가 어디까지 사실인지 검증합니다.]]></summary></entry><entry xml:lang="ko"><title type="html">머스크의 다보스 예측을 다시 읽는다: 5년 뒤의 지능, 그리고 그 지능은 어디서 도는가</title><link href="https://thakicloud.github.io/ko/culture/three-year-window-labor-to-assets/" rel="alternate" type="text/html" title="머스크의 다보스 예측을 다시 읽는다: 5년 뒤의 지능, 그리고 그 지능은 어디서 도는가" /><published>2026-07-20T00:00:00+09:00</published><updated>2026-07-20T00:00:00+09:00</updated><id>https://thakicloud.github.io/ko/culture/three-year-window-labor-to-assets</id><content type="html" xml:base="https://thakicloud.github.io/ko/culture/three-year-window-labor-to-assets/"><![CDATA[<p><img src="/assets/images/three-year-window-labor-to-assets-hero.png" alt="소멸하는 인간의 시간이 휴머노이드와 데이터센터의 빛으로 흘러 들어가는 추상 이미지" /></p>

<h2 id="개요">개요</h2>

<p>최근 한 게시글이 여러 커뮤니티를 돌았습니다. 일론 머스크에게 말을 거는 형식이었고, 요지는 강렬했습니다. 사람이 노동으로 돈을 벌 수 있는 시간이 길게 잡아도 앞으로 3년 정도이며, 그 뒤에는 지금 급여를 주던 일자리 대부분이 자동화되어 사람에게 돈을 지불할 이유가 사라진다는 것이었습니다. 글은 이것을 역사상 최대 규모의 부의 이동이라 부르며, 지금 해야 할 일은 자기 시간을 기계에 빼앗기지 않는 자산으로 바꾸는 것이라고 결론지었습니다.</p>

<p>그런데 그 게시글에는 텍스트만 있는 것이 아니라 실제 영상이 붙어 있었습니다. 머스크 본인이 말하는 인터뷰 클립이었습니다. 그래서 저희는 원 영상을 받아 그가 정확히 무엇을 말했는지 한 문장씩 확인해 보았습니다. 확인해 보니 흥미로운 사실이 드러났습니다. 화제가 된 “3년”이라는 숫자는 머스크가 말한 것이 아니었고, 정작 그가 말한 숫자는 따로 있었습니다. 이 글에서는 먼저 그가 실제로 한 말을 정리하고, 그 예측을 진지하게 받아들일 때 어떤 그림이 그려지는지, 그리고 기술 블로그다운 각도에서 그 지능이라는 것이 대체 어디서 돌아가는지를 짚겠습니다.</p>

<!-- Courtesy of embedresponsively.com -->

<div class="responsive-video-container">
    <iframe src="https://drive.google.com/file/d/1HczL43lXw-P-geWxtPEHQc_eiapkDPwx/preview" frameborder="0" webkitallowfullscreen="" mozallowfullscreen="" allowfullscreen=""></iframe>
  </div>

<p>위 영상이 화제가 된 클립입니다. 세계경제포럼(다보스)에서 블랙록 CEO 래리 핑크와 나눈 대담의 일부로 알려져 있습니다.</p>

<h2 id="머스크가-실제로-한-말">머스크가 실제로 한 말</h2>

<p>영상을 받아 적어 보면 머스크의 발언은 세 개의 예측으로 압축됩니다. 원문을 그대로 옮기면 이렇습니다. “앞으로 5년, 그러니까 2031년쯤이면 디지털 지능이 모든 인간 지능의 총합을 넘어설 것이라고 봅니다.” “5년 안에 휴머노이드 로봇이 적어도 1억 대, 어쩌면 10억 대에 이를 수 있습니다.” “경제 규모는 5년, 길게는 6~7년 안에 지금의 두 배가 될 것이라고 예측합니다. 배증 주기에 들어서기 때문입니다. 경제 산출이 워낙 빨리 늘어서, 몇 년의 오차를 감안하더라도 거대한 변화를 보게 될 것입니다.”</p>

<p>정리하면 세 가지입니다. 첫째, 약 5년 뒤 디지털 지능이 인류 전체 지능의 합을 넘어선다. 둘째, 같은 기간에 휴머노이드 로봇이 1억에서 10억 대 규모로 늘어난다. 셋째, 경제가 5~7년 안에 두 배로 커진다. 세 예측 모두 시간 단위가 3년이 아니라 5년입니다.</p>

<p>여기서 정확성을 한 번 짚고 가겠습니다. 화제가 된 게시글의 “3년”은 머스크의 발언이 아니라, 그 클립을 인용한 작성자가 자기 해석을 얹어 만든 표현입니다. 머스크는 5년 뒤 지능과 로봇과 경제의 변화를 말했고, 작성자는 그 변화의 앞자락, 즉 사람이 노동으로 버틸 수 있는 시간을 3년으로 압축해 극적으로 표현한 것입니다. 두 숫자를 구별하는 것이 이 논의의 출발점입니다. 원본에 충실하자면 우리가 다뤄야 할 시한은 3년이 아니라 5년 안팎이고, 다룰 주제는 자산 종목이 아니라 지능과 노동의 관계입니다.</p>

<h2 id="왜-이-예측을-진지하게-받아들이는가">왜 이 예측을 진지하게 받아들이는가</h2>

<p>시한부 예언은 대개 조심스럽게 다뤄야 하지만, 이번 예측의 방향에는 무시하기 어려운 근거가 있습니다. 저희도 이 예측의 큰 방향에는 대체로 동의하는 편이며, 그 이유를 근거와 함께 밝혀 두겠습니다.</p>

<p>첫째, 지능 쪽입니다. 지난 몇 해 사이 언어 모델이 코드를 쓰고, 문서를 작성하고, 고객 문의를 처리하고, 데이터를 분석하는 능력은 많은 이의 예상보다 빠르게 올라왔습니다. “이건 아직 못 하겠지” 하던 일들이 계속 넘어갔습니다. 학습에 투입되는 연산과 데이터가 커질수록 성능이 예측 가능하게 개선되는 흐름이 이어지는 한, 특정 시점에 집합적 지능의 저울이 기계 쪽으로 기우는 그림은 공상이 아니라 추세의 연장선입니다.</p>

<p>둘째, 로봇 쪽입니다. 지능이 소프트웨어에 머물지 않고 몸을 얻어 물리 세계의 노동으로 확장되는 것이 휴머노이드입니다. 대수가 1억이냐 10억이냐는 제조 역량과 비용 곡선에 달린 문제이고 오차가 크겠지만, 방향 자체, 즉 물리 노동의 상당 부분이 기계로 넘어가기 시작한다는 방향은 이미 여러 공장과 물류 현장에서 조짐이 보입니다.</p>

<p>셋째, 경제 쪽입니다. 지능과 노동의 공급이 급격히 늘면 산출이 늘고, 산출이 빠르게 늘면 경제는 배증 주기에 들어섭니다. 머스크의 5~7년 배증 예측은 낙관적이지만, 그 메커니즘 자체는 경제학이 오래 다뤄 온 논리입니다. 시점이 몇 년 어긋나더라도 방향은 성립합니다.</p>

<p>정직하게 덧붙이자면, 방향이 옳다는 것과 시점이 정확하다는 것은 다른 문제입니다. 기술 예측에는 로이 아마라의 오래된 관찰이 잘 들어맞습니다. 사람은 기술의 단기 효과를 과대평가하고 장기 효과를 과소평가한다는 말입니다. 5년이라는 숫자는 어긋날 수 있습니다. 그러나 그 방향까지 어긋난다고 보기는 어렵습니다. 방향은 맞고 속도는 불확실하다고 보는 편이 정직합니다.</p>

<h2 id="노동에서-자산으로-그-논리의-단단한-부분과-무른-부분">노동에서 자산으로, 그 논리의 단단한 부분과 무른 부분</h2>

<p>원 게시글의 결론은 급여를 “자기 시간을 빌려주는 대가”로 정의하는 데서 출발합니다. 그 시간이 하는 일을 기계가 더 싸고 안정적으로 해낸다면, 빌려주던 시간의 값이 떨어지고 결국 무엇을 소유했는가만 남는다는 것입니다.</p>

<p>단단한 부분부터 봅시다. 경제에서 만들어지는 소득이 노동과 자본으로 나뉜다고 할 때, 자동화는 그 저울을 자본 쪽으로 기울입니다. 기계가 사람의 일을 대신할수록 그 기계를 소유한 쪽으로 이익이 더 돌아가는 것은 산업혁명 이래 반복된 패턴이고, 이번의 AI와 로봇은 그 패턴을 지식노동과 물리노동 양쪽으로 넓히고 있습니다. “무엇을 소유했는가가 중요해진다”는 방향 자체는 무리한 주장이 아닙니다.</p>

<p>무른 부분은 그다음입니다. 첫째, 일자리는 통째로 사라지기보다 모양을 바꾸는 쪽이 역사적으로 더 흔했습니다. 자동화는 특정 과업을 없애면서 동시에 그것을 설계하고 감독하고 붙여 쓰는 새로운 과업을 만들어 냈습니다. 다만 이번에는 새로 생기는 과업마저 빠르게 자동화될 수 있다는 점에서, 과거의 낙관을 그대로 반복하기는 어렵습니다. 둘째, “그러니 특정 자산을 지금 사라”는 처방에는 비약이 있습니다. 자본으로 소득이 기운다는 진단과, 어떤 자산을 지금 사야 한다는 처방 사이에는 큰 간격이 있습니다. 저희는 투자 자문을 제공하지 않으며, 특정 자산의 가격은 변동이 크다는 점을 분명히 해 둡니다. 이 글이 다루는 것은 종목이 아니라 그 아래 깔린 구조입니다.</p>

<h2 id="그-지능은-어디서-도는가">그 지능은 어디서 도는가</h2>

<p>여기서 기술 블로그다운 각도로 한 걸음 들어가겠습니다. 머스크가 말한 디지털 지능도, 휴머노이드 로봇의 두뇌도, 경제를 두 배로 밀어 올린다는 산출도, 전부 허공에서 일어나지 않습니다. 모델을 학습시키고 추론을 돌리고 로봇을 제어하는 실제 연산이 어딘가의 GPU 위에서 일어납니다. 다시 말해 “인류 지능의 총합을 넘어서는 지능”의 물리적 실체는 컴퓨트, 그리고 그 컴퓨트를 돌리는 전력과 인프라입니다.</p>

<p>골드러시에서 진짜 돈을 번 쪽이 금을 캔 사람보다 곡괭이와 청바지를 판 사람이었다는 오래된 비유가 여기 그대로 들어맞습니다. 지능과 노동이 자동화될수록, 그 자동화가 소비하는 연산과 전력의 값은 올라갑니다. 그런데 여기서 원 게시글의 자산론과 갈라지는 중요한 차이가 하나 생깁니다. 어떤 자산은 소유하되 아무것도 생산하지 않지만, 컴퓨트 인프라는 소유하는 동시에 실제 일을 만들어 냅니다. 전자가 가치를 저장하는 그릇이라면, 후자는 가치를 생성하는 공장입니다. 자동화 시대에 “무엇을 소유할 것인가”라는 질문의 가장 생산적인 답은, 자동화 그 자체를 돌리는 능력을 소유하는 것입니다.</p>

<div class="mermaid">
flowchart TB
    A["사람의 노동시간<br />(급여로 교환)"] --&gt;|자동화로 대체·보완| B["기계가 수행하는 일<br />(지식노동 + 물리노동)"]
    B --&gt; C{"만들어진 가치는<br />어디에 쌓이나?"}
    C --&gt;|가치를 저장| D["희소 자산<br />(생산하지 않는 그릇)"]
    C --&gt;|가치를 생성| E["컴퓨트·전력·인프라<br />(지능을 돌리는 공장)"]
    E --&gt; F["누가 그 인프라를<br />소유·통제하는가"]
    F --&gt; G["개인: 대체되지 않는 역량<br />조직: 자기 컴퓨트와 데이터"]
</div>

<p>개인 차원으로 옮기면 이 통찰은 조금 덜 극적이지만 더 실천적입니다. 자산을 서둘러 사들이는 것보다, 자동화를 설계하고 지휘하고 검증하는 능력, 즉 기계가 쉽게 대체하지 못하는 판단과 맥락을 자기 안에 쌓는 편이 대부분의 사람에게 더 확실한 대비입니다. 도구를 두려워하는 대신 도구를 부리는 자리로 옮겨 앉는 것이야말로, 시간을 자산으로 바꾸는 가장 정직한 방법입니다.</p>

<h2 id="thakicloud의-렌즈로-본다면">ThakiCloud의 렌즈로 본다면</h2>

<p>저희가 이 주제를 남의 이야기로 다루지 않는 이유가 있습니다. ThakiCloud가 만드는 것이 정확히 그 “지능이 도는 밑바닥”이기 때문입니다.</p>

<p>첫째는 ai-platform 렌즈입니다. ThakiCloud의 ai-platform은 쿠버네티스 위에서 GPU 자원을 배분하고, 모델을 학습시키고, 추론을 서빙하는 AI/ML 인프라입니다. 여기서 저희가 특히 무게를 두는 지점이 있습니다. 컴퓨트를 남의 클라우드에 통째로 맡기는 대신, 조직이 자기 환경에서 자기 모델과 자기 데이터를 직접 돌릴 수 있게 하는 것입니다. 온프렘과 소버린 AI, 낮은 서빙 비용, 셀프 호스팅이 저희가 반복해서 붙드는 키워드입니다. 앞의 논리를 그대로 옮기면, 지능이 소비하는 연산을 스스로 소유하고 통제하는 능력을 조직에게 돌려주는 일입니다. 데이터 주권이 중요한 공공과 규제 산업일수록 이 각도의 무게가 큽니다.</p>

<p>둘째는 Paxis 렌즈입니다. 컴퓨트를 소유하는 것만으로는 절반입니다. 그 컴퓨트를 실제 자동화된 일로 바꾸는 제어 평면이 있어야 합니다. Paxis는 ai-platform 위에서 도는 Agent-Native Cloud로, 스킬과 툴과 정책과 감사 로그를 일급 자원으로 다룹니다. 수백 개의 스킬을 선택해 격리된 샌드박스에서 실행하고, 여러 에이전트를 DAG로 엮어 협업시키며, 모든 행동을 정책 게이트와 감사 로그로 통과시킵니다. 앞서 컴퓨트를 가치를 생성하는 공장이라 불렀는데, Paxis는 그 공장의 작업 지시서이자 안전장치에 해당합니다. 저비용 서빙이 자동화의 경제성을 만들고, 그 위의 에이전트 계층이 그 경제성을 실제 업무 성과로 바꾸는 구조입니다.</p>

<p>이 두 렌즈를 합치면 머스크의 예측이 불러일으키는 불안에 대한 저희 나름의 답이 됩니다. 자동화가 피할 수 없는 방향이라면, 그 흐름의 통제권을 소수의 손에 넘기지 않고 더 많은 조직이 나눠 갖도록 만드는 것, 그것이 인프라 회사가 할 수 있는 몫이라고 저희는 봅니다.</p>

<h2 id="신중하게-볼-지점">신중하게 볼 지점</h2>

<p>큰 방향에 동의하면서도, 이 예측을 다룰 때 흔들리지 않기 위해 붙들어야 할 지점이 몇 가지 있습니다.</p>

<p>먼저 시점입니다. 5년이라는 숫자는 어긋날 수 있고, 화제가 된 3년은 더 그렇습니다. 완전 자율주행이 매년 “내년이면 된다”고 불린 지 십 년이 넘었다는 사실을 기억하면, 정확한 연도를 못박는 예언일수록 그 연도보다 그것이 불러일으키는 감정이 목적일 때가 많습니다. 방향을 믿되 달력을 믿지는 않는 태도가 필요합니다.</p>

<p>다음은 처방입니다. 자본으로 소득이 기운다는 진단이 옳다는 것과, 지금 무엇을 사야 한다는 것은 전혀 다른 이야기입니다. 후자는 대개 검증되지 않은 채 확신의 어조로 전달됩니다. 확신에 찬 문장일수록 근거를 따로 떼어 확인하는 습관이 필요합니다.</p>

<p>마지막은 분배입니다. 부의 이동이 실제로 크다면, 그것은 개인의 영리한 자산 선택만으로 풀릴 문제가 아니라 사회가 함께 다뤄야 할 문제입니다. 자산을 살 수 있는 사람만을 향한 처방은, 애초에 그 처방이 진단한 문제를 더 키울 위험이 있습니다. 머스크 자신도 다른 자리에서 보편적 고소득 같은 개념을 언급한 바 있는데, 이는 자동화 이후의 분배가 개인 차원이 아니라 제도 차원의 질문임을 시사합니다.</p>

<h2 id="맺음">맺음</h2>

<p>이 클립에서 확실하게 건질 수 있는 것은 하나입니다. 자동화는 방향으로서 실재하고, 그 가치는 자동화를 실제로 돌리는 능력, 즉 개인에게는 대체되지 않는 역량으로, 조직에게는 자기 컴퓨트와 데이터로 쌓인다는 것입니다. 머스크가 말한 5년이 정확한 달력이 아니더라도, 지능과 노동의 관계가 바뀌고 있다는 방향에는 저희도 대체로 동의합니다. 다만 그 문 너머에 있는 것은 서둘러 사야 할 종목이 아니라 천천히 쌓아 올려야 할 능력과 인프라라고, 저희는 고쳐 읽습니다.</p>

<h2 id="출처">출처</h2>

<ul>
  <li>원 영상: 일론 머스크가 세계경제포럼(다보스)에서 블랙록 CEO 래리 핑크와 나눈 대담 클립. 본문의 발언은 이 영상을 직접 받아 적어 확인한 것입니다.</li>
  <li>발언 요지 교차 확인: 머스크의 5년 예측(디지털 지능이 인류 지능 총합을 넘어섬, 휴머노이드 로봇 1억~10억 대, 경제 5~7년 내 배증)은 다보스 발언으로 여러 매체가 보도했습니다.</li>
  <li>화제가 된 “3년” 표현은 해당 클립을 인용한 게시글 작성자의 해석이며, 머스크 본인의 발언이 아닙니다.</li>
  <li>로이 아마라의 관찰(단기 과대평가·장기 과소평가)은 기술 예측을 다룰 때 널리 인용되는 경험칙입니다.</li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="culture" /><category term="AI자동화" /><category term="미래의노동" /><category term="컴퓨트인프라" /><category term="소버린AI" /><category term="에이전트경제" /><category term="일론머스크" /><category term="기술철학" /><summary type="html"><![CDATA[일론 머스크가 다보스에서 던진 세 가지 예측이 '3년 뒤 급여의 시대가 끝난다'는 경고로 번졌습니다. 실제 인터뷰 영상을 받아 적어 그가 정확히 무엇을 말했는지 확인하고, 노동에서 자산으로 넘어가는 흐름의 밑바닥에서 도는 컴퓨트 인프라가 왜 진짜 관건인지 짚습니다.]]></summary></entry></feed>