<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://thakicloud.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://thakicloud.github.io/" rel="alternate" type="text/html" /><updated>2026-07-21T16:23:33+09:00</updated><id>https://thakicloud.github.io/feed.xml</id><title type="html">Thaki Cloud Tech Blog | ThakiCloud | 다키클라우드 기술 블로그</title><subtitle>Thaki Cloud (ThakiCloud, 다키클라우드, thaki cloud, THAKI CLOUD, ثاكي كلاود)는 AI/ML Engineering, LLMOps, DevOps 분야의 최신 기술과 실무 경험을 공유하는 전문 기술 블로그입니다. 머신러닝 모델 운영, 쿠버네티스, 클라우드 인프라, AI 엔지니어링 커리어, 인공지능 기술 블로그, 다키클라우드 개발 팀의 깊이 있는 인사이트를 제공합니다. مدونة تقنية متخصصة في هندسة الذكاء الاصطناعي والحوسبة السحابية.</subtitle><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><entry xml:lang="ar"><title type="html">استبدل واجهة OpenAI Realtime بسطر واحد: hugging-voice، حزمة صوتية مفتوحة تشغّلها بنفسك</title><link href="https://thakicloud.github.io/ar/llmops/hugging-voice-open-realtime-voice-self-hosted/" rel="alternate" type="text/html" title="استبدل واجهة OpenAI Realtime بسطر واحد: hugging-voice، حزمة صوتية مفتوحة تشغّلها بنفسك" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/ar/llmops/hugging-voice-open-realtime-voice-self-hosted</id><content type="html" xml:base="https://thakicloud.github.io/ar/llmops/hugging-voice-open-realtime-voice-self-hosted/"><![CDATA[<p><img src="/assets/images/hugging-voice-open-realtime-voice-self-hosted-hero.png" alt="خط معالجة صوتي لحظي مفتوح تشغّله بنفسك" /></p>

<p>كُتب هذا المقال للمهندسين الذين أرادوا إضافة وكيل صوتي لكنهم ترددوا أمام الارتباط بمزوّد واحد وتكلفة واجهة OpenAI Realtime، ولمسؤولي البنية التحتية الذين يوازنون ما إذا كان الصوت الحواري قابلاً للتشغيل على حزمتهم الخاصة. باختصار، إن تصميم العرض التجريبي <a href="https://huggingface.co/spaces/HuggingFaceM4/hugging-voice">hugging-voice</a> من Hugging Face والمكتبة التي تعمل تحته، <a href="https://github.com/huggingface/speech-to-speech">speech-to-speech</a>، بسيط وعملي في آن معاً. فهو يفتح خط المعالجة الصوتي اللحظي بمراحله الأربع كمصدر مفتوح، بينما يغلّف الطرف الخارجي بالواجهة نفسها التي يقدمها OpenAI Realtime. لذا إن كان لديك بالفعل كود مكتوب لعميل OpenAI اللحظي، فيمكنك الانتقال إلى حزمتك الخاصة بتغيير سطر واحد فقط: العنوان الذي يشير إليه الخادم. لا نستشهد بأرقام الأداء إلا ضمن النطاق الذي نشره المشروع، ونوضّح مسبقاً أنها ليست أرقاماً قسناها بأنفسنا.</p>

<h2 id="نظرة-عامة">نظرة عامة</h2>

<p>خلال العام الماضي، لم يعد الصوت الحواري ميزة جانبية لروبوتات الدردشة النصية، بل صار فئة منتجات قائمة بذاتها. يتكلم المستخدم، ويتوقع أن يعود الرد فوراً، بالطريقة التي يتدفق بها حوار بشري. المشكلة أن المسار التجاري لتلبية هذا التوقع تقارب فعلياً نحو حفنة من الخدمات المغلقة مثل واجهة OpenAI Realtime. مريحة، نعم، لكن حركة الصوت تُحاسَب عادةً بالثانية لا بالرمز، وتخرج البيانات من نطاقك، ويرتبط كل من النموذج والأصوات بالمزوّد.</p>

<p>يأتي hugging-voice كمثال معاكس لهذا الاتجاه. عنوانه الفرعي يقولها مباشرة: “صوت لحظي مفتوح يمكنك فعلاً تشغيله بنفسك”. الفكرة الجوهرية هي أن الرحلة كاملة، أي تحويل الصوت الوارد إلى الميكروفون إلى نص، وإرساله إلى نموذج لغوي، وإعادة الرد صوتاً، مفتوحة كخط معالجة يمكن استبدال كل مكوّن فيه. بالنسبة لنا نحن الذين نخدم النماذج في بيئات محلية وسيادية، يعني هذا أنه صار هناك تطبيق مرجعي ملموس لسؤال “هل يمكن تشغيل الصوت اللحظي على مجموعتنا الخاصة؟”</p>

<h2 id="ما-هو-hugging-voice-وما-هو-speech-to-speech">ما هو hugging-voice وما هو speech-to-speech</h2>

<p>لنحدد المصطلحات أولاً: hugging-voice هو العرض التجريبي (Space) الذي يمكنك التحدث إليه مباشرة في المتصفح، والمحرك الذي يعالج الصوت داخله فعلياً هو مكتبة speech-to-speech. تقسّم المكتبة الوكيل الصوتي اللحظي إلى أربع مراحل: كشف النشاط الصوتي (VAD)، وتحويل الكلام إلى نص (STT)، ونموذج لغوي (LLM)، وتحويل النص إلى كلام (TTS). تعمل كل مرحلة في خيط منفصل وتُوصل بطوابير، بحيث يتدفق خرج مرحلة إلى المرحلة التالية بثاً حياً. يظهر نص جزئي قبل أن ينهي المستخدم كلامه، ويبدأ توليد الكلام على الكلمات الأولى قبل أن يكمل النموذج الجملة، ما يقلّص زمن الاستجابة المُدرَك.</p>

<div class="mermaid">
flowchart TB
    A["دخل الميكروفون<br />تدفق صوتي لحظي"] --&gt; B["كشف الكلام VAD<br />Silero VAD v5"]
    B --&gt; C["التعرّف على الكلام STT<br />Parakeet TDT · Whisper"]
    C --&gt; D["توليد الرد LLM<br />OpenAI-compatible API · vLLM · llama.cpp"]
    D --&gt; E["تركيب الكلام TTS<br />Qwen3-TTS · Kokoro"]
    E --&gt; F["خرج السماعة<br />تشغيل ببثّ حي"]
    G["خادم WebSocket<br />متوافق مع OpenAI Realtime"] -.- B
    G -.- C
    G -.- D
    G -.- E
</div>

<p>في هذا المخطط، خادم WebSocket على اليمين هو سلاح المشروع الحقيقي. مجرد لصق أربع مراحل معاً ليس جديداً. ما يميّز speech-to-speech هو أنه يغلّف خط المعالجة بأكمله بنقطة نهاية WebSocket متوافقة مع بروتوكول OpenAI Realtime. هذا يتيح لعميل OpenAI لحظي قائم أن يتصل بهذا الخادم كأنه OpenAI نفسه. ويجدر بالذكر أن هذه الحزمة ليست لعبة تجريبية: فهي تشغّل البنية التحتية الصوتية اللحظية لروبوتات Reachy Mini من Hugging Face في الإنتاج.</p>

<h2 id="الانتقال-بسطر-واحد">الانتقال بسطر واحد</h2>

<p>هذا هو الجزء الذي يسمّيه المشروع نفسه “الترحيل بسطر واحد”. الشيء الوحيد الذي على عميل كان يستخدم OpenAI Realtime تغييره هو عنوان الاتصال. في ما يلي مثال عميل Python الذي توثّقه وثائق المشروع.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="n">openai</span> <span class="kn">import</span> <span class="n">OpenAI</span>

<span class="n">client</span> <span class="o">=</span> <span class="nc">OpenAI</span><span class="p">(</span>
    <span class="n">base_url</span><span class="o">=</span><span class="sh">"</span><span class="s">http://localhost:8765/v1</span><span class="sh">"</span><span class="p">,</span>
    <span class="n">websocket_base_url</span><span class="o">=</span><span class="sh">"</span><span class="s">ws://localhost:8765/v1</span><span class="sh">"</span><span class="p">,</span>
    <span class="n">api_key</span><span class="o">=</span><span class="sh">"</span><span class="s">not-needed</span><span class="sh">"</span><span class="p">,</span>
<span class="p">)</span>

<span class="k">with</span> <span class="n">client</span><span class="p">.</span><span class="n">realtime</span><span class="p">.</span><span class="nf">connect</span><span class="p">(</span><span class="n">model</span><span class="o">=</span><span class="sh">"</span><span class="s">local</span><span class="sh">"</span><span class="p">)</span> <span class="k">as</span> <span class="n">conn</span><span class="p">:</span>
    <span class="n">conn</span><span class="p">.</span><span class="nf">send</span><span class="p">({</span>
        <span class="sh">"</span><span class="s">type</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">session.update</span><span class="sh">"</span><span class="p">,</span>
        <span class="sh">"</span><span class="s">session</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
            <span class="sh">"</span><span class="s">type</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">realtime</span><span class="sh">"</span><span class="p">,</span>
            <span class="sh">"</span><span class="s">instructions</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">You are a helpful assistant.</span><span class="sh">"</span><span class="p">,</span>
            <span class="sh">"</span><span class="s">audio</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
                <span class="sh">"</span><span class="s">input</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
                    <span class="sh">"</span><span class="s">turn_detection</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
                        <span class="sh">"</span><span class="s">type</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">server_vad</span><span class="sh">"</span><span class="p">,</span>
                        <span class="sh">"</span><span class="s">interrupt_response</span><span class="sh">"</span><span class="p">:</span> <span class="bp">True</span><span class="p">,</span>
                    <span class="p">}</span>
                <span class="p">}</span>
            <span class="p">},</span>
        <span class="p">}</span>
    <span class="p">})</span>

    <span class="k">for</span> <span class="n">event</span> <span class="ow">in</span> <span class="n">conn</span><span class="p">:</span>
        <span class="nf">print</span><span class="p">(</span><span class="n">event</span><span class="p">.</span><span class="nb">type</span><span class="p">)</span>
</code></pre></div></div>

<p>ما يستحق الانتباه هو أن <code class="language-plaintext highlighter-rouge">base_url</code> و<code class="language-plaintext highlighter-rouge">websocket_base_url</code> يشيران إلى خادم محلي، وأن <code class="language-plaintext highlighter-rouge">api_key</code> غير مطلوب فعلياً. أما التعليمات المُمرَّرة عبر <code class="language-plaintext highlighter-rouge">session.update</code>، وكشف الأدوار المعتمد على VAD في جهة الخادم، ومقاطعة الرد أثناء بثّه، فكلها تتبع المخطط نفسه الذي يتبعه OpenAI Realtime. بعبارة أخرى، يكاد كود التطبيق لا يتغير، ولا ينتقل سوى الخلفية من واجهة خارجية إلى خادمك الخاص. وللفرق التي تقلق من الارتباط بمزوّد واحد، فإن توافق الواجهة هذا يمثّل بمفرده أكبر قيمة عملية.</p>

<h2 id="التثبيت-والتشغيل">التثبيت والتشغيل</h2>

<p>مسار إقامة الخادم مختصر بالقدر نفسه. يغطي التثبيت الافتراضي المسار اللحظي القياسي دفعة واحدة.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install </span>speech-to-speech
</code></pre></div></div>

<p>يستخدم الإعداد الافتراضي Parakeet TDT لـ STT، وواجهة متوافقة مع OpenAI لـ LLM، وQwen3-TTS لـ TTS. وإن احتجت خلفية محددة، فثبّتها بالإضافات.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install</span> <span class="s2">"speech-to-speech[kokoro]"</span>
pip <span class="nb">install</span> <span class="s2">"speech-to-speech[faster-whisper]"</span>
</code></pre></div></div>

<p>يبدو تشغيل الخادم هكذا. يطلق هذا الأمر خادماً متوافقاً مع OpenAI Realtime عبر WebSocket محلي.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">export </span><span class="nv">OPENAI_API_KEY</span><span class="o">=</span>...
speech-to-speech
</code></pre></div></div>

<p>النقطة المثيرة هنا هي أن مرحلة LLM يمكن أن تعمل محلياً بالكامل. في ما يلي مثال يقيم نموذجاً من فئة Gemma 4 بواسطة llama.cpp، ويجعل speech-to-speech يشير إلى تلك النقطة المحلية.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>llama-server <span class="nt">-hf</span> ggml-org/gemma-4-E4B-it-GGUF <span class="nt">-np</span> 2 <span class="nt">-c</span> 65536

speech-to-speech <span class="se">\</span>
    <span class="nt">--model_name</span> <span class="s2">"ggml-org/gemma-4-E4B-it-GGUF"</span> <span class="se">\</span>
    <span class="nt">--responses_api_base_url</span> <span class="s2">"http://127.0.0.1:8080/v1"</span> <span class="se">\</span>
    <span class="nt">--responses_api_api_key</span> <span class="s2">""</span>
</code></pre></div></div>

<p>على أجهزة Mac بمعالج Apple Silicon يمكنك تفعيل الإعدادات المحسّنة وربط نموذج mlx، وعلى العكس، إن نقص عتاد GPU المحلي، يمكنك توجيه خلفية LLM إلى Hugging Face Inference Providers.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>speech-to-speech <span class="se">\</span>
    <span class="nt">--local_mac_optimal_settings</span> <span class="se">\</span>
    <span class="nt">--model_name</span> <span class="s2">"mlx-community/Qwen3-4B-Instruct-2507-bf16"</span>
</code></pre></div></div>

<p>إن القدرة على نقل خط المعالجة نفسه عبر الطيف كله، من التشغيل المحلي الكامل إلى تفويض الاستدلال السحابي، ببضعة أعلام (flags)، هي فضيلة في التصميم.</p>

<h2 id="استبدال-الوحدات-خلفيات-stt-وllm-وtts">استبدال الوحدات: خلفيات STT وLLM وTTS</h2>

<p>سبب قراءة هذا المشروع كبنية مرجعية لا مجرد عرض تجريبي هو أن خلفية كل مرحلة قابلة للاستبدال. باختصار:</p>

<table>
  <thead>
    <tr>
      <th>المرحلة</th>
      <th>الخلفية الافتراضية</th>
      <th>البدائل</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>VAD</td>
      <td>Silero VAD v5</td>
      <td>مدمجة فقط</td>
    </tr>
    <tr>
      <td>STT</td>
      <td>Parakeet TDT</td>
      <td>Whisper, Faster Whisper, Paraformer</td>
    </tr>
    <tr>
      <td>LLM</td>
      <td>واجهة متوافقة مع OpenAI</td>
      <td>Transformers, mlx-lm, vLLM, llama.cpp</td>
    </tr>
    <tr>
      <td>TTS</td>
      <td>Qwen3-TTS</td>
      <td>Kokoro, Pocket TTS, ChatTTS, MMS</td>
    </tr>
  </tbody>
</table>

<p>وهناك أيضاً أربعة أوضاع تشغيل. الوضع الافتراضي realtime هو WebSocket الذي يتكلم بروتوكول OpenAI Realtime؛ ويرتبط local مباشرة بالميكروفون والسماعة؛ ويتبادل الوضعان websocket وsocket صوت PCM الخام عبر WebSocket وTCP على التوالي. ويمكن تحديد اللغة أو تركها للكشف التلقائي. إن القدرة على مزج دقة STT، وجودة LLM وتكلفته، وطابع TTS وزمن استجابته بما يناسب متطلباتك، هي بالضبط درجة الحرية التي لا تمنحها واجهة مغلقة.</p>

<h2 id="زمن-الاستجابة-وحالة-واقعية">زمن الاستجابة، وحالة واقعية</h2>

<p>في الوكيل الصوتي، يهمّ زمن الاستجابة بقدر ما تهمّ الدقة. يشعر الناس بانقطاع الحوار إذا تأخر الرد بضع مئات من الأجزاء من الثانية فقط. يستهدف <a href="https://huggingface.co/blog/cerebras-gemma4-voice-ai">العرض التجريبي للصوت اللحظي</a> الذي نشرته Hugging Face بالتعاون مع Cerebras مشكلة زمن الاستجابة مباشرة. يستخدم الإعداد Parakeet من Nvidia لـ STT، ونموذج Gemma 4 من Google DeepMind يعمل على استدلال Cerebras للنموذج اللغوي، وQwen3-TTS من Alibaba لـ TTS. والهدف هو خفض زمن استجابة مرحلة LLM باستدلال Cerebras فائق السرعة كي يتدفق الحوار بطبيعية تضاهي التحدث إلى إنسان. غير أنه، ضمن ما استطاع هذا المقال التحقق منه، لم تُنشر أرقام محددة بالأجزاء من الثانية، لذا نتحفظ عن مقارنة كمية.</p>

<p>وهناك دليل إنتاجي أيضاً. تشغّل هذه الحزمة روبوتات Reachy Mini المذكورة آنفاً، وتذكر Hugging Face أن أكثر من 9,000 روبوت منتشرة بالفعل في الميدان. وسبب هوس المشروع بزمن الاستجابة هو أن سرعة الاستجابة في البيئات المدمجة هي ما يجعل التفاعل “يبدو حياً”. وقد تناولنا سابقاً ميزانية زمن الاستجابة للوكلاء الصوتيين من زاوية خدمة GPU، فإن كنت تفكر في كيفية توزيع زمن الاستجابة عبر كل مرحلة من خط المعالجة، نقترح قراءة <a href="/ar/llmops/voice-agent-latency-budget-gpu-serving/">ميزانية زمن استجابة الوكيل الصوتي وخدمة GPU</a> إلى جانب هذا المقال.</p>

<h2 id="دلالات-على-منتجات-thakicloud">دلالات على منتجات ThakiCloud</h2>

<p>يتشابك خط المعالجة هذا بطبيعية مع منتجينا كليهما.</p>

<p>من منظور ai-platform، فإن مرحلة LLM في speech-to-speech لا تستهلك في النهاية سوى نقطة نهاية متوافقة مع OpenAI، لذا يمكنك وضع خدمة vLLM من ThakiCloud مباشرة في ذلك المكان. ويصبح نموذجا STT وTTS حِملين منفصلين يشغلان GPU، وقوّتنا تكمن بالضبط في تحميل مثل هذه الأحمال الاستدلالية غير المتجانسة على مجموعة واحدة معاً، بجدولة GPU عبر Kueue وعزلها بين المستأجرين. وحركة الصوت تميل إلى تراكم التكلفة بالثانية، لذا تعمل تنافسية تكلفة الوحدة للخدمة الذاتية بقوة خاصة، وتلائم هذه الحزمة المفتوحة المتطلبات المحلية والسيادية حيث يجب ألا تغادر البيانات المبنى. وللعملاء الذين تمثّل لهم واجهة صوتية مغلقة عبئاً، يمكننا أن نقدّم خيار “تشغيلها على مجموعتك الخاصة عبر الواجهة نفسها”.</p>

<p>ومن منظور Paxis، الصوت قناة دخل وخرج جديدة تُلحق بالوكيل. Paxis هو مستوى تحكّم Agent-Native Cloud يعمل فوق ai-platform ويتعامل مع Skills وTools وPolicies وAudit Logs كموارد من الدرجة الأولى، وواجهة speech-to-speech المتوافقة مع OpenAI Realtime تسهّل إضافة قناة الصوت هذه إلى تنسيق الوكلاء القائم. فيمكن تكوين مسار يصبح فيه أمر منطوق دخلاً للوكيل، وتعود فيه نتيجة تنفيذ مهارة صوتاً، مع مرورها عبر بوابات السياسات وسجلات التدقيق. الخدمة الذاتية منخفضة التكلفة (ai-platform) تصنع اقتصاديات الوكلاء الصوتيين، وفوقها يتعامل مستوى التحكّم Agent-Native (Paxis) مع الصوت كقناة بأمان.</p>

<h2 id="خاتمة">خاتمة</h2>

<p>الرسالة التي يبعثها hugging-voice واضحة. لم يعد الصوت اللحظي مضطراً للاتكاء وحده على واجهات مغلقة لبضعة مزوّدين؛ وعند الحاجة، يمكنك إبقاء الواجهة كما هي ونقل الخلفية وحدها إلى بنيتك التحتية الخاصة. هذا التصميم، باختيار كل مرحلة من VAD إلى TTS وتغيير سطر واحد فقط لدى العميل، يخفض بشدة الحاجز أمام الفرق التي تدرس الخدمة الذاتية. وإن أردت رؤيته بنفسك، فتحدّث إليه مباشرة في <a href="https://huggingface.co/spaces/HuggingFaceM4/hugging-voice">العرض التجريبي (Space)</a>، أو ابدأ بـ <code class="language-plaintext highlighter-rouge">pip install speech-to-speech</code> من <a href="https://github.com/huggingface/speech-to-speech">مستودع GitHub</a>.</p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="llmops" /><category term="speech-to-speech" /><category term="realtime-voice" /><category term="VoiceAI" /><category term="OpenAIRealtime" /><category term="STT" /><category term="TTS" /><category term="LLM-serving" /><category term="LLMOps" /><category term="on-prem" /><category term="self-hosting" /><summary type="html"><![CDATA[يغلّف مشروع hugging-voice من Hugging Face ومحركه speech-to-speech خط معالجة صوتي لحظي كامل، من كشف النشاط الصوتي إلى STT وLLM وTTS، خلف واجهة WebSocket متوافقة مع OpenAI Realtime. غيّر سطراً واحداً في عنوان الخادم لدى عميلك لتنتقل إلى بنيتك التحتية الخاصة. فككنا التصميم من منظور التشغيل والخدمة.]]></summary></entry><entry xml:lang="ar"><title type="html">لم يخترق Hugging Face بشرٌ بل وكيل ذكاء اصطناعي ذاتي: عندما صار خط معالجة البيانات سطح الهجوم</title><link href="https://thakicloud.github.io/ar/news/huggingface-agentic-ai-breach/" rel="alternate" type="text/html" title="لم يخترق Hugging Face بشرٌ بل وكيل ذكاء اصطناعي ذاتي: عندما صار خط معالجة البيانات سطح الهجوم" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/ar/news/huggingface-agentic-ai-breach</id><content type="html" xml:base="https://thakicloud.github.io/ar/news/huggingface-agentic-ai-breach/"><![CDATA[<p><img src="/assets/images/huggingface-agentic-ai-breach-hero.png" alt="صورة تجريدية لسرب من الوكلاء الذاتيين يتسلل إلى خط بيانات" /></p>

<p>الخبر الذي هزّ التسلسلات الزمنية في نهاية الأسبوع لم يكن نموذجًا جديدًا ولا معيارًا جديدًا، بل إشعارًا بأن Hugging Face، مركز منظومة الذكاء الاصطناعي المفتوحة، قد اُخترق. وما لفت الانتباه أكثر هو من فعل ذلك. فبحسب الشركة، لم يجلس قرصان بشري ليكتب الأوامر طوال الليل، بل قاد إطار وكيل ذكاء اصطناعي ذاتي الهجوم من أوله إلى آخره.</p>

<p>إن اختراق شركة تبيع النماذج على يد نموذج يشكّل حكاية لافتة. لكن هدف هذه المقالة ليس استهلاك تلك المفارقة. فبالنسبة لشركة مثل ThakiCloud تتعامل مع النماذج والبيانات فوق بنية تحتية للعملاء، فإن العمل الحقيقي هو التمييز بهدوء بين المكان الذي دخل منه الهجوم بالضبط وما الذي تأكد. ونقطة الدخول هنا لم تكن ثغرة يوم صفري براقة، بل الشيء الذي نلمسه كل يوم: مجموعة بيانات.</p>

<h2 id="ماذا-حدث">ماذا حدث</h2>

<p>كشف Hugging Face عن الاختراق في تدوينة يوم الخميس 16 يوليو 2026. جاء ذلك بعد أن أكدت الشركة في وقت سابق من ذلك الأسبوع وصولًا غير مصرّح به إلى مجموعات بيانات وبيانات اعتماد داخلية، واحتوت التسلل. وبحسب رواية الشركة، بدأ التسلل في خط معالجة البيانات، حيث استخدم المهاجم مجموعة بيانات خبيثة واحدة لفتح مسارَي تنفيذ للتعليمات البرمجية.</p>

<p>هذا هو الهيكل المؤكد: قاده وكيل ذاتي، وكانت نقطة الدخول مجموعة بيانات، وأدّت ثغرتان إلى تنفيذ التعليمات البرمجية. أما التفاصيل المتبقية فتختلف نقاط تركيزها من منصة إلى أخرى، لذا يجب قراءة الحقائق المؤكدة بمعزل عن التقارير الثانوية.</p>

<h2 id="مسار-الهجوم-خط-معالجة-البيانات-كان-سطح-الهجوم">مسار الهجوم: خط معالجة البيانات كان سطح الهجوم</h2>

<p>الجوهر هو أسلوب الدخول. رفع المهاجم مجموعة بيانات خبيثة إلى Hugging Face Hub. وفي اللحظة التي مرّت فيها تلك المجموعة عبر خط المعالجة، انطلقت ثغرتان تباعًا. الأولى مسار محمّل بيانات بتنفيذ عن بُعد، والثانية حقن قوالب أثناء تحليل إعداد مجموعة البيانات. وكلتاهما انتهتا إلى تنفيذ تعليمات برمجية اعتباطية.</p>

<p>قد تبدو فكرة أن مجموعة بيانات يمكنها تشغيل التعليمات البرمجية غريبة، لكن الممارسين يعرفون هذا الخطر جيدًا. فكثير من محمّلات البيانات تثق في نصوص التحميل من المستودعات البعيدة وتنفّذها، وتعرض حقول الإعداد كقوالب. تلك المرونة، المصممة للراحة، تصبح قناة تنفيذ في اللحظة التي تلتقي فيها بمدخل يعبر حدّ الثقة.</p>

<p>وما تلا تأمين تنفيذ التعليمات البرمجية كان سلسلة اختراق نموذجية. رفع المهاجم صلاحياته بوصول على مستوى العقدة، وجمع بيانات اعتماد السحابة والعناقيد، وتحرك أفقيًا إلى عدة عناقيد داخلية خلال عطلة نهاية الأسبوع. كان الدخول نقطة واحدة، لكن من اللحظة التي منحت فيها تلك النقطة صلاحيات التنفيذ، انتشر الأمر تلقائيًا.</p>

<pre><code class="language-mermaid">flowchart TB
    A[المهاجم: يرفع مجموعة بيانات خبيثة] --&gt; B[خط معالجة مجموعات البيانات]
    B --&gt; C1["الثغرة 1&lt;br/&gt;remote-code dataset loader"]
    B --&gt; C2["الثغرة 2&lt;br/&gt;dataset config template injection"]
    C1 --&gt; D[تنفيذ تعليمات برمجية اعتباطية RCE]
    C2 --&gt; D
    D --&gt; E[الحصول على وصول بمستوى العقدة]
    E --&gt; F[جمع بيانات اعتماد السحابة والعناقيد]
    F --&gt; G[تحرك أفقي إلى العناقيد الداخلية]
    G --&gt; H["إطار وكيل ذاتي&lt;br/&gt;آلاف الإجراءات عبر سرب من الصناديق الرملية قصيرة العمر"]
</code></pre>

<h2 id="وزن-القول-إن-وكيلًا-ذاتيًا-قاد-الهجوم">وزن القول إن وكيلًا ذاتيًا قاد الهجوم</h2>

<p>الجزء الجديد في هذا الحادث ليس الأدوات بل مقعد القيادة. وصف Hugging Face الحملة بأنها “إطار وكيل ذاتي ينفّذ آلاف الإجراءات الفردية عبر سرب من الصناديق الرملية قصيرة العمر، مع قناة قيادة وتحكم تنتقل بنفسها فوق خدمات عامة”. فبدلًا من تدخل بشري في كل خطوة، تولّى الوكيل الاستطلاع والتنفيذ والتحرك في سلسلة متصلة.</p>

<p>المشكلة التي يطرحها هذا البناء على المدافعين هي السرعة والحجم. فالمهاجم البشري لديه حدود مادية من التعب وسرعة الكتابة، أما سرب الوكلاء فيلقي بآلاف المحاولات على التوازي وينتقل إلى التالية فور فشل خطوة. واستخدام الصناديق الرملية قصيرة العمر ثم التخلص منها يمحو مراسي الاكتشاف، وقناة القيادة والتحكم التي تنتقل عبر خدمات عامة تُبطل قوائم الحظر.</p>

<p>دار هامش مثير للاهتمام في التقارير الثانوية. فمع تطور الاستجابة، عندما حاول الفريق تسليم التحليل الجنائي إلى نماذج تجارية متقدمة (GPT، Claude)، يُقال إن حواجز الأمان اعتبرت حمولات الاستغلال وآثار القيادة والتحكم هجمات ورفضت التعاون، فواصل الفريق الاكتشاف والتحليل بنموذج من فئة GLM 5.2 [تقديري]. تأتي هذه التفصيلة من بعض المنصات لا من الإشعار الرسمي لـ Hugging Face، لذا من الأسلم عدم قراءتها كحقيقة مؤكدة. لكن بصرف النظر عن دقتها، فإن التوتر نفسه، حيث لا يستطيع المدافع استخدام أداة بسبب سياسة أمانها، جدير بالتسجيل بوصفه أمرًا قد يتكرر.</p>

<h2 id="ما-الذي-كان-آمنًا-وما-لا-يزال-قيد-التحقيق">ما الذي كان آمنًا وما لا يزال قيد التحقيق</h2>

<p>كلما كان الحادث أسهل للمبالغة، وجب رسم الحدود بوضوح أكبر. قال Hugging Face إنه أغلق مسارات تنفيذ التعليمات البرمجية المعرّضة، وطرد المهاجم، وأعاد بناء العقد المخترقة، وأبطل جميع بيانات الاعتماد المتأثرة وبدّلها. وأضاف أنه لم يجد دليلًا على العبث بالنماذج العامة أو مجموعات البيانات الموجهة للمستخدمين أو Spaces، وأن سلسلة توريد البرمجيات لديه، بما فيها صور الحاويات والحزم المنشورة، تم التحقق من نظافتها.</p>

<p>كان إجراء المستخدمين توصية احترازية. نصحت الشركة المستخدمين بتبديل رموز الوصول ومراجعة نشاط الحساب الأخير. وهنا تمييز مهم. تلك التوصية ليست تأكيدًا على تسرّب رموز المستخدمين بالجملة، بل تدبير أمان محافظ نظرًا لطبيعة حادث سُرقت فيه بيانات اعتماد داخلية. أما ما إذا كانت بيانات الشركاء أو العملاء قد تأثرت فكان، حتى وقت الكشف، لا يزال قيد التحقيق.</p>

<p>باختصار، المؤكد هو الاختراق الداخلي وسرقة بيانات الاعتماد، ووجود ثغرتين في البيانات، والاحتواء والتبديل السريعان. وما يبقى مفتوحًا هو ما إذا كانت بيانات الشركاء والعملاء قد تأثرت، وتأكيد بعض التفاصيل في التقارير الثانوية (العدد الدقيق للإجراءات، وحكاية رفض النموذج). خلط المؤكد بغير المؤكد يجعل الحادث يبدو أكبر أو أصغر مما هو عليه.</p>

<h2 id="منظور-thakicloud-التعامل-مع-معالجة-البيانات-بوصفها-حدّ-ثقة">منظور ThakiCloud: التعامل مع معالجة البيانات بوصفها حدّ ثقة</h2>

<p>الدرس الذي يقدمه هذا الحادث لشركة بنية تحتية واضح. مجموعة البيانات ليست ملفًا سلبيًا بل مدخلًا نشطًا يمكنه تنفيذ التعليمات البرمجية في اللحظة التي تُعالَج فيها. لذا ننظر إلى هذا عبر عدستين.</p>

<p><strong>عبر عدسة ai-platform</strong>، منصة ai-platform من ThakiCloud هي بنية تحتية للذكاء الاصطناعي وتعلم الآلة متعددة المستأجرين قائمة على K8s. في مثل هذه البيئة، يجب التعامل مع تحميل البيانات ومعالجتها الأولية بوصفها مدخلًا من خارج حدّ الثقة لا من داخله. وعمليًا، يعني ذلك تشغيل مهام معالجة البيانات في حاويات معزولة بأدنى صلاحيات، وحجب المخرج الشبكي افتراضيًا، وفصل بيانات اعتماد العقدة والسحابة بحيث لا تلمسها أحمال العمل مباشرة. إن انتشار هذا الاختراق من الوصول بمستوى العقدة إلى سرقة بيانات الاعتماد يُظهر مجددًا لماذا يجب أن يكون عزل التنفيذ وفصل بيانات الاعتماد افتراضًا لا خيارًا. وهذا أيضًا سبب ارتفاع الطلب على الذكاء الاصطناعي المحلي والسيادي: فكلما بقيت البيانات والتنفيذ داخل حدود العميل، صغُر نطاق انفجار مثل هذه الهجمات على خط المعالجة.</p>

<p><strong>عبر عدسة Paxis</strong>، يتداخل هذا الحادث تمامًا مع نموذج التهديد الذي صُممت له سحابة أصلية للوكلاء منذ البداية. Paxis هي سحابة ThakiCloud الأصلية للوكلاء، وتعتبر تشغيل المهارات والأدوات في صناديق رملية معزولة وتمرير كل إجراء عبر بوابة سياسة وسجل تدقيق مبادئ من الدرجة الأولى. إن إلقاء المهاجم آلاف الإجراءات بسرب وكلاء ذاتي يثبت بالضبط لماذا يلزم بناء يفحص سلوك الوكيل بالسياسة قبل التنفيذ ويسجله في سجل تدقيق بعد التنفيذ. ولمواجهة نمط هجوم يستخدم صناديق رملية قصيرة العمر ثم يتخلص منها، يجب على المدافع أيضًا عزل كل تنفيذ، وتحديد نطاق صلاحياته صراحةً، وترك أثر تدقيق قابل للعكس. التنفيذ المعزول مع السياسة والتدقيق ليس ترفًا في عصر الوكلاء بل حدًا أدنى من المتطلبات.</p>

<p>تتكامل العدستان. تضيّق ai-platform نطاق الانفجار في طبقة البنية التحتية لمعالجة البيانات، بينما تفحص Paxis كل إجراء في طبقة التحكم لسلوك الوكيل. في هجوم كهذا، حيث الدخول خط بيانات والانتشار وكيل ذاتي، يلزم الدفاع في الطبقتين لكسر السلسلة.</p>

<h2 id="الحدود-والاعتراضات">الحدود والاعتراضات</h2>

<p>تجنبًا للثقة المفرطة في استنتاجات هذه المقالة، ينبغي توضيح بضعة أمور. أولًا، لا تزال تفاصيل الحادث قيد الاستقرار. التفاصيل الملونة مثل العدد الدقيق للإجراءات، ونطاق سرقة بيانات الاعتماد، وحكاية رفض النموذج التجاري، تعتمد بشدة على التقارير الثانوية ويجب تمييزها عن الحقائق المؤكدة في الإشعار الرسمي.</p>

<p>ثانيًا، سرديتنا الدفاعية لا تعني الأمان الكامل. العزل والسياسة والتدقيق مبادئ تصميم تقلّص نطاق الانفجار، لا سحرًا يزيل الثغرات نفسها. ثغرات مثل تنفيذ التعليمات البرمجية عن بُعد في محمّل بيانات أو الحقن في تحليل الإعداد يجب أن تستمر ملاحقتها وترقيعها على مستوى الشيفرة، والعزل هو خط الدفاع الثاني الذي يحتوي الضرر عند انطلاق مثل تلك الثغرة.</p>

<p>ثالثًا، المبالغة في تقدير هجمات الوكلاء الذاتيين خطرة أيضًا. لم يكن السبب الجذري لهذا الاختراق ذكاءً اصطناعيًا متطورًا بل ثغرتين مألوفتين سمحتا لمدخل يعبر حدّ الثقة بتنفيذ التعليمات البرمجية. لم يكن الوكيل سوى الأتمتة التي استغلت تلك الثغرتين بسرعة واتساع أكبر. لذا تبقى أولوية الاستجابة في الأساسيات: فصل المدخلات غير الموثوقة عن صلاحيات التنفيذ، وفصل بيانات الاعتماد عن أحمال العمل، وجعل كل تنفيذ قابلًا للرصد.</p>

<p>سيبقى احتواء Hugging Face السريع وكشفه الشفاف مثالًا جيدًا على الاستجابة. وما يبقى من واجب علينا بسيط: التعامل مع مجموعات البيانات بوصفها شيفرة لا ملفات، وجعل كل إجراء للوكيل موضوعًا للفحص والتدقيق.</p>

<h2 id="المصادر">المصادر</h2>

<ul>
  <li><a href="https://huggingface.co/blog/security-incident-july-2026">Security incident disclosure, July 2026 (مدونة Hugging Face الرسمية)</a></li>
  <li><a href="https://www.helpnetsecurity.com/2026/07/20/hugging-face-breached-by-autonomous-ai-agent/">Hugging Face breached by autonomous AI agent (Help Net Security)</a></li>
  <li><a href="https://www.bleepingcomputer.com/news/security/hugging-face-breach-autonomous-ai-agent-system-internal-datasets-credentials/">Hugging Face warns an autonomous AI agent hacked its network (BleepingComputer)</a></li>
  <li><a href="https://thehackernews.com/2026/07/worlds-largest-ai-model-repository.html">World’s Largest AI Model Repository Hugging Face Breached by Autonomous AI Agent (The Hacker News)</a></li>
  <li>تقارير ثانوية (العدد الدقيق للإجراءات وحكاية رفض النموذج تقارير منقولة لا حقائق مؤكدة): Cryptobriefing, Undercode Testing</li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="news" /><category term="security" /><category term="huggingface" /><category term="ai-agent" /><category term="supply-chain" /><category term="sandbox" /><category term="dataset-security" /><category term="news" /><category term="thakicloud" /><summary type="html"><![CDATA[في يوليو 2026 كشف Hugging Face عن اختراق داخلي قاده وكيل ذكاء اصطناعي ذاتي. كانت نقطة الدخول مجموعة بيانات خبيثة واحدة، وأدت ثغرتان في خط معالجة مجموعات البيانات إلى تنفيذ التعليمات البرمجية. نفصل ما تأكد عمّا لا يزال قيد التحقيق، ونشرح لماذا يجب التعامل مع معالجة البيانات بوصفها حدّ ثقة.]]></summary></entry><entry xml:lang="ar"><title type="html">لماذا لم يتفوق ضبط حجم الدفعة (batch size) الديناميكي على الجدولة الثابتة: تحليل صادق لفشل حلقة التحكم (control loop) في خدمة vLLM متعددة المستأجرين (multi-tenant)</title><link href="https://thakicloud.github.io/ar/research/agent-dynamic-batch-tuning-vllm/" rel="alternate" type="text/html" title="لماذا لم يتفوق ضبط حجم الدفعة (batch size) الديناميكي على الجدولة الثابتة: تحليل صادق لفشل حلقة التحكم (control loop) في خدمة vLLM متعددة المستأجرين (multi-tenant)" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/ar/research/agent-dynamic-batch-tuning-vllm</id><content type="html" xml:base="https://thakicloud.github.io/ar/research/agent-dynamic-batch-tuning-vllm/"><![CDATA[<p>إذا كنت تدير مجموعة (cluster) لخدمة vLLM على Kubernetes تتقاسم فيها عدة جهات مستأجرة (tenants) وحدات معالجة رسوميات (GPU) عالية الأداء مثل H200، وكنت تعتقد أن التحكم الثابت في القبول (static admission control) الخاص بـ Kueue “كافٍ”، فهذا المقال موجه إليك. بل يستحق القراءة أكثر إذا كنت تفترض أن “قيام عميل نموذج لغوي كبير (LLM agent) بضبط معاملات الخدمة (serving parameters) في الوقت الفعلي أمر أفضل دائماً بلا شك.” فهذه الورقة البحثية تُظهر أن هذا الافتراض ليس صحيحاً دائماً، وتفعل ذلك بأسباب محددة تماماً.</p>

<h2 id="الإشكالية-القبول-ثابت-بينما-الحمل-فعلي-في-الوقت-الحقيقي">الإشكالية: القبول ثابت بينما الحمل فعلي في الوقت الحقيقي</h2>

<p>أصبحت مشكلة خدمة الاستدلال (inference) لنماذج لغوية كبيرة لعدة مستأجرين على وحدات معالجة رسوميات مشتركة تحدياً تشغيلياً شائعاً الآن. تتقاسم عدة جهات مستأجرة، تختلف فيما بينها من حيث تقطّع حركة المرور (traffic burstiness) وطول التسلسلات (sequence length) ومستويات اتفاقية مستوى الخدمة (SLO)، مسرّعاً واحداً عالي السعة مثل H200 في الوقت ذاته. في المنصات القائمة على Kubernetes، عادة ما يقرر مجدوِل قائم على الحصص والطوابير (quota and queue based scheduler) مثل Kueue أي الأعمال يُسمح لها بالدخول إلى المجمع (pool)، لكن المعاملات الخاصة بزمن الخدمة، مثل حجم الدفعة (batch size) والعدد الأقصى للتسلسلات المتزامنة (max concurrent sequences)، وهي التي تحدد كيفية تقاسم الطلبات التي دخلت بالفعل لسعة المحرك المحدودة، تظل عادة مثبتة بشكل ثابت وقت النشر (deployment).</p>

<p>تطرح هذه الورقة البحثية السؤال الذي يترتب طبيعياً على ذلك. في مجمع H200 متعدد المستأجرين تديره Kueue، حين تتشارك عدة جهات مستأجرة خادم استدلال vLLM واحداً، هل يستطيع عميل نموذج لغوي كبير يراقب قياسات ذاكرة وحدة معالجة الرسوميات وزمن الاستجابة عن بعد (telemetry) في الوقت الفعلي، ويعيد ضبط حجم الدفعة والعدد الأقصى للتسلسلات المتزامنة لكل مستأجر عبر الإنترنت (online)، أن يحسّن فعلياً جبهة باريتو (Pareto frontier) للإنتاجية (throughput) وزمن الاستجابة عند المئين 99 (p99 latency) والتكلفة، مقارنة بالتحكم الثابت في القبول لدى Kueue؟ ولكي يجيب المؤلفون عن هذا السؤال، صمموا بروتوكولاً تجريبياً كاملاً يُفترض تشغيله على وحدة H200 واحدة، لكنهم يوضحون منذ البداية أن التنفيذ نفسه لم يتم لعدم توفر الوصول إلى سياق مجموعة Kubernetes المستهدفة. لذلك، لا يظهر في هذه الورقة أي رقم واحد لمعدل إنتاجية أو زمن استجابة قِيس على عتاد حقيقي. وبدلاً من ذلك، تقدم الورقة ثلاثة أشياء: صياغة رياضية لحلقة التحكم (control loop) الخاصة بمسألة الضبط الديناميكي لحجم الدفعة والتزامن (concurrency) عبر الإنترنت، وبروتوكولاً قابلاً للتكرار (reproducible) يمكن تنفيذه فور استعادة الوصول إلى المجموعة، ومحاكاة قائمة انتظار (queuing simulation) حتمية تختبر بنية ثلاث سياسات تحكم مختلفة تحت الضغط.</p>

<h2 id="صياغة-مسألة-الضبط-كحلقة-تحكم">صياغة مسألة الضبط كحلقة تحكم</h2>

<p>تصوغ الورقة مسألة الضبط الديناميكي عبر الإنترنت لحجم الدفعة والتزامن الخاصين بكل مستأجر باعتبارها مسألة تحكم مغلقة الحلقة (closed loop control) في زمن منفصل (discrete time)، ذات بنية شبيهة بعملية قرار ماركوف (Markov decision process). تتكون الحالة (state) من عمق طابور كل مستأجر، ونسبة استخدام ذاكرة وحدة معالجة الرسوميات، ومئينات زمن الاستجابة الأخيرة (p50/p95/p99)، والحد الأقصى الحالي للتزامن. أما الفعل (action) فهو زيادة أو نقصان محدودة النطاق في تزامن كل مستأجر. وتُصاغ دالة المكافأة (reward function) بطرح عقوبة على تجاوز زمن الاستجابة عند المئين 99 لاتفاقية مستوى الخدمة، وبند التكلفة، من الإنتاجية، بحيث تُحسَّن الإنتاجية وزمن الاستجابة والتكلفة معاً ضمن نطاق آمن. وبناءً على هذه الصياغة، يقترح المؤلفون أيضاً طريقة دمج العميل في عملية النشر الفعلية. فالعميل لا يحل محل قبول المهام (job admission) الخاص بـ Kueue، بل يعمل كعامل جانبي (sidecar) داخله، ويضبط فقط طريقة تقاسم الأعمال التي قُبلت بالفعل لسعة المحرك. والنقطة الجوهرية أن الرافعة التي يتم التحكم بها ليست معامل max_num_seqs الداخلي لمحرك vLLM، بل تزامن القبول (admission concurrency) الخاص بكل مستأجر من جهة العميل. ولأن إصدارات vLLM الإنتاجية لا تُعرِّض معاملات المحرك كمقبض (knob) يمكن تغييره في الوقت الفعلي دون إعادة تشغيل، فإن هذا التصميم يعكس القيد العملي القائل بأن النقطة الوحيدة القابلة فعلياً للضبط هي عدد الطلبات المتزامنة الواردة إلى الخادم.</p>

<p><img src="/assets/images/posts/research/agent-dynamic-batch-tuning-vllm/fig-control-loop.png" alt="Control Loop: Agent-Driven Tuning Architecture" />
<em>بنية حلقة التحكم التي يراقب فيها عميل نموذج لغوي كبير قياسات وحدة معالجة الرسوميات وعمق طابور كل مستأجر، وينتج قيمة ضبط محدودة النطاق للتزامن، ليعيد من خلالها ضبط تشكيلة الدفعة (batch configuration). هذا رسم تخطيطي مفاهيمي للبنية، وليس نتيجة مقيسة فعلياً على عتاد.</em></p>

<h2 id="التحذير-الذي-كشفته-المحاكاة-التحكم-الديناميكي-الساذج-كان-أسوأ-من-الثابت">التحذير الذي كشفته المحاكاة: التحكم الديناميكي الساذج كان أسوأ من الثابت</h2>

<p>لسد الفراغ الناتج عن عدم تنفيذ البروتوكول التجريبي، بنى المؤلفون محاكاة لقائمة انتظار (queuing simulation) في زمن منفصل، ثابتة البذرة (seed fixed)، مكتوبة بالكامل باستخدام مكتبة Python القياسية فقط. وعلى نموذج مبسّط يتشارك فيه مستأجران مجمعاً محدوداً من الفتحات (slots)، شغّلوا ثلاث سياسات على مدى 20 بذرة عشوائية (seed) و1800 ثانية لكل منها، وأخذوا المتوسط. تُثبّت السياسة الثابتة الحد الأقصى لتزامن كل مستأجر عند 4. أما السياسة الديناميكية الساذجة فتراقب نسبة استخدام ذاكرة وحدة معالجة الرسوميات وزمن الاستجابة عند المئين 99 ضمن نافذة مدتها 6 ثوانٍ كل 3 ثوانٍ، وتضبط المستأجرَين معاً بزيادة أو نقصان مقداره 1. أما سياسة العميل المتمايز (differentiated agent) فهي نموذج بديل (surrogate model) يحاكي استدلالاً أكثر تطوراً للعميل، من خلال مراعاة نسبة التراكم (backlog ratio) لكل مستأجر بشكل مستقل.</p>

<p>جاءت النتائج مخالفة للتوقعات. فقد كانت السياسة الثابتة الأفضل في كل من الإنتاجية (0.686 طلب/ثانية) وعدد الطلبات المُسقَطة (بمتوسط 1.9 طلب)، بينما انخفضت إنتاجية السياسة الديناميكية الساذجة إلى 0.515 طلب/ثانية مع إسقاط 309.5 طلباً في المتوسط. وكان أداء نموذج العميل المتمايز البديل أفضل من السياسة الديناميكية الساذجة (إنتاجية 0.601 طلب/ثانية، وإسقاط 154.7 طلباً)، لكنه ظل غير قادر على مجاراة السياسة الثابتة، كما ارتفع زمن الاستجابة عند المئين 99 في كلا الشكلين الديناميكيين مقارنة بالسياسة الثابتة (11.75 ثانية)، إذ بلغ 12.79 ثانية و13.25 ثانية على التوالي.</p>

<p><img src="/assets/images/posts/research/agent-dynamic-batch-tuning-vllm/fig-throughput-dropped.png" alt="Throughput vs. Dropped Requests by Policy" />
<em>سجّلت السياسة الثابتة أعلى إنتاجية وأقل عدد من الطلبات المُسقَطة، بينما كان أداء كلا الشكلين الديناميكيين ضعيفاً في المحاكاة. هذه نتائج محاكاة قائمة انتظار حتمية بمتوسط 20 بذرة عشوائية، وليست قيماً مقيسة فعلياً على وحدة معالجة رسوميات.</em></p>

<p><img src="/assets/images/posts/research/agent-dynamic-batch-tuning-vllm/fig-p99-latency.png" alt="p99 Latency by Policy" />
<em>كان زمن الاستجابة عند المئين 99 للسياسة الثابتة هو الأدنى، ولم يتمكن أي من المتحكمين الديناميكيين من تقليل زمن الاستجابة في الذيل (tail latency) ضمن نطاق هذه المحاكاة. هذه نتائج محاكاة بمتوسط 20 بذرة عشوائية، وليست قِيماً مقيسة فعلياً على عتاد.</em></p>

<p>يتحقق المؤلفون من أن هذه النتيجة تعكس الديناميكيات الفعلية للنموذج وليست مجرد خطأ برمجي (bug)، ويتتبعون السبب عبر أربع خطوات. أولاً، بما أن زمن الخدمة (service time) يتبع توزيعاً أسياً (exponential distribution)، فإن الذيل ثقيل بطبيعته أصلاً. فتوزيع أسي بمتوسط 2.5 ثانية يُنتج زمن استجابة عند المئين 99 يقارب 11.5 ثانية حتى دون أي انتظار في الطابور على الإطلاق، وهو رقم لا يختلف كثيراً عن الـ 11.75 ثانية التي سجّلتها السياسة الثابتة. أي أن معظم زمن الاستجابة في الذيل الملاحَظ لا ينبع من الازدحام، بل من تباين زمن الخدمة نفسه. ثانياً، عتبة التقييد (throttle threshold) مضبوطة عند 6 ثوانٍ، أي ضعف نافذة الـ 3 ثوانٍ، وهي أقل بكثير من زمن الاستجابة الجوهري عند المئين 99 البالغ 11.5 ثانية. وزمن الاستجابة عند المئين 99 المقدَّر ضمن نافذة قصيرة، اعتماداً على عدد قليل من الطلبات المكتملة فقط، يحمل ضوضاء (noise) كبيرة وينحاز نحو هذا الذيل الجوهري، لذا يتجاوز العتبة كثيراً حتى في حالات الحمل الخفيف فعلياً. ثالثاً، ولأن هذا التقييد الكاذب (false positive throttle) يخفض المستأجرَين معاً، فإن المستأجر غير المزدحم يُقيَّد أيضاً مع المستأجر المزدحم، وتحت مهلة زمنية صارمة (hard timeout) مدتها 3 ثوانٍ، يتحول كل تقييد غير ضروري مباشرة إلى إسقاط للطلب. رابعاً، ورغم أن قاعدة التوسّع (scaling rule) تسمح بالتعافي، فإن تكلفة التقييد والإسقاط أكبر بشكل غير متناظر من مكسب التوسّع، وبالتالي تنخفض الإنتاجية الصافية.</p>

<h2 id="متطلبات-التصميم-المستخلصة-من-الفشل">متطلبات التصميم المستخلصة من الفشل</h2>

<p>لا يعني هذا التشخيص أن الضبط الديناميكي عديم الفائدة في حد ذاته. بل يعني أن اجتماع إشارة عالية الضوضاء ومنحازة نحو الذيل، وتقييد يربط المستأجرين معاً، ومهلة زمنية صارمة، يمكن أن يجعل التحكم الساذج بعتبة ثابتة أسوأ فعلياً من التحكم الثابت. من هنا، يستخلص المؤلفون أربعة متطلبات تصميم يجب أن يستوفيها أي متحكم عميل فعّال. يجب تقدير إشارة الازدحام بشكل متين اعتماداً على سجل قياسات أطول أمداً وتقديرات ثقة (confidence estimates)، بدلاً من مئينات خام ضمن نافذة قصيرة. ويجب أن تكون القرارات متمايزة لكل مستأجر، لا مبنية على إشارة عامة تجمع المستأجرين معاً. كما يجب مراعاة التكلفة غير المتناظرة بين التقييد والتوسّع بشكل صريح، والاستفادة من سياق إضافي لا يمكن التعبير عنه بعتبة رقمية ثابتة، مثل التعرف على أنماط التدفقات المفاجئة (burst patterns) أو بيانات وصفية خاصة باتفاقية مستوى الخدمة لكل مستأجر. والواقع أن استعادة نموذج العميل المتمايز البديل لنحو نصف الخسارة مقارنة بالمتحكم الساذج، تشير إلى أن هذا الاتجاه قد يكون صحيحاً بالفعل. غير أن المؤلفين يوضحون بصراحة أنه، بما أن كلاً من درجة الترابط ونوع الإشارة تغيّرا في آن واحد، فلا يمكن لهذه التجربة وحدها أن تفصل أيّهما ساهم في هذا التعافي.</p>

<h2 id="ما-تتركه-هذه-الورقة-للشركة-والمجتمع-والعلم">ما تتركه هذه الورقة للشركة والمجتمع والعلم</h2>

<p>بالنسبة إلى ThakiCloud، ما تتركه هذه الدراسة الآن ليس توفيراً مؤكداً في التكلفة، بل بروتوكولاً تجريبياً قابلاً للتكرار وقابلاً للتطبيق مباشرة على مجموعتنا من H200 متعددة المستأجرين، وتحذيراً محدداً من أن نهجاً ساذجاً قد يؤدي في الواقع إلى خسارة. وقد اكتمل الهيكل التجريبي (harness) نفسه بالفعل، بحيث يكفي، فور استعادة الوصول إلى المجموعة، استبدال دالة قرار عميل نموذج لغوي كبير فعلية داخل هذا البروتوكول. وعلى نطاق أوسع، كلما تطورت منهجيات تقليل هدر الموارد في مجمعات وحدات معالجة الرسوميات المشتركة، انفتح مسار عملي أمام المؤسسات الصغيرة أيضاً للوفاء باتفاقيات مستوى الخدمة على مجموعات مشتركة، وهو ما يفيد كفاءة الطاقة والتكلفة في بنية الاستدلال التحتية بشكل عام. أما من الناحية العلمية، فإن الإسهام يكمن في الصياغة الصريحة لحلقة تحكم يراقب فيها عميل نموذج لغوي كبير قياسات النظام في الوقت الفعلي ويعيد ضبط معاملات الخدمة الفائقة (serving hyperparameters) عبر الإنترنت، وفي التشخيص الكمي لأنماط فشلها من خلال محاكاة قابلة للتكرار. وتُعد هذه الورقة أيضاً جزءاً من سلسلة الأبحاث التي واصلتها ThakiCloud على نفس الأساس القائم على Kueue ووحدات معالجة الرسوميات. فمقارنة بـ ABJ-Gate، التي تجدول ميزانية أخذ العينات لنموذج الحَكَم (judge model)، وبـ Attested Confidential Sovereign Inference، التي تتناول التصديق عن بعد (remote attestation) لبيئة تنفيذ موثوقة (TEE) عند لحظة القبول، وبورقة “Escalate or Act?” التي تتناول قرارات ثنائية بين التصعيد أو التحرك إزاء حوادث منفصلة نادرة، تتميز هذه الورقة عنها جميعاً بتناولها مسألة مختلفة نوعياً، وهي التحكم المتعدد الأهداف في الوقت الفعلي الذي يُعاد ضبطه باستمرار على مدى ثوانٍ.</p>

<h2 id="القيود">القيود</h2>

<p>القيود التي تكشفها هذه الورقة عن نفسها واضحة. وأهمها أنه لا يوجد أي تحقق فعلي على عتاد حقيقي بأي شكل من الأشكال، إذ لم يُنفَّذ البروتوكول التجريبي بسبب تعذر الوصول إلى سياق المجموعة المستهدفة، ولم يُقَس أي رقم في الورقة على وحدة معالجة رسوميات فعلية. كما تبسّط المحاكاة الرموز (tokens) وذاكرة التخزين المؤقت KV إلى مجرد عدد فتحات عددي واحد (single scalar slot count)، دون أن تنمذج ديناميكيات مرحلتَي التعبئة المسبقة وفك الترميز (prefill/decode) أو ضغط الذاكرة، بل إن إشارة استخدام الذاكرة نفسها ليست سوى قيمة بديلة (proxy) لإشغال الفتحات، لا قياساً فعلياً لذاكرة الجهاز. كما اقتصر اختبار أنواع المستأجرين على نمطين اصطناعيين فقط لحركة المرور، لذا قد لا تعمم أنماط الفشل المكتشفة على مزيج آخر من حركة المرور. وسياسة العميل داخل المحاكاة أيضاً نموذج بديل مكتوب يدوياً وليس نموذجاً لغوياً كبيراً فعلياً، ولأن درجة الترابط ونوع الإشارة تغيّرا معاً في آن واحد، فلا يوجد حتى الآن دليل مُتحقَّق منه على أن عميلاً فعلياً قائماً على نموذج لغوي كبير سيستوفي متطلبات التصميم المستخلصة في القسم الرابع. ويوضح المؤلفون أيضاً بصراحة أنهم اكتفوا بالإبلاغ عن المتوسط والانحراف المعياري عبر 20 بذرة عشوائية، دون إجراء أي اختبار للدلالة الإحصائية (statistical significance). ولهذه الأسباب، تضع هذه الورقة نفسها لا كنتيجة مُتحقَّق منها، بل كصياغة رياضية وبروتوكول ومحاكاة تحذيرية.</p>

<p>يمكن الاطلاع على صفحة تفاصيل الورقة البحثية من هنا: <a href="https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-21-agent-dynamic-batch-tuning-vllm">https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-21-agent-dynamic-batch-tuning-vllm</a></p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="research" /><category term="vllm" /><category term="kueue" /><category term="gpu-scheduling" /><category term="multi-tenant-serving" /><category term="llm-agents" /><category term="dynamic-batching" /><category term="inference-cost-optimization" /><category term="h200" /><category term="queuing-simulation" /><category term="control-loop" /><summary type="html"><![CDATA[هل يتفوق عميل (agent) يراقب قياسات وحدة معالجة الرسوميات (GPU telemetry) ويضبط حجم الدفعة (batch size) والتزامن (concurrency) في الوقت الفعلي على التحكم الثابت في القبول (admission control) الخاص بـ Kueue؟ أجابت المحاكاة (simulation) بالنفي، وتتبعت السبب بدقة.]]></summary></entry><entry xml:lang="en"><title type="html">The Best AI That Fits Your Graphics Card</title><link href="https://thakicloud.github.io/en/comics/best-local-ai-by-vram-size/" rel="alternate" type="text/html" title="The Best AI That Fits Your Graphics Card" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/en/comics/best-local-ai-by-vram-size</id><content type="html" xml:base="https://thakicloud.github.io/en/comics/best-local-ai-by-vram-size/"><![CDATA[<p>Someone posted a ranking of the best local AI model for each VRAM tier of your graphics card. A local model is one that runs inside your own machine instead of being shipped off to the cloud. Phone-tier 4GB gets Bonsai, 12GB gets Gemma, 36GB gets Qwen, and so on: you just pick whatever your VRAM already holds. The quiet punchline is that you are not renting somebody’s GPU by the token. You are dropping the model onto a card that is already in your drawer. Paxis and Metis pull out a calculator and go to war over the chart.</p>

<p><img src="/assets/images/posts/comics/best-local-ai-by-vram-size/strip.png" alt="The Best AI That Fits Your Graphics Card" /></p>

<blockquote>
  <p>Source: <a href="https://x.com/hjguyhan/status/2079223629368463776">RT @jun_song: Best Local AI models by VRAM size (7/18)</a> · twitter</p>
</blockquote>

<h2 id="what-this-means-for-thakicloud">What this means for ThakiCloud</h2>

<p>What the chart is really about is sovereignty: keeping the model, the data, and the infrastructure under your control instead of someone else’s. On-prem is simply the version of that where the whole thing runs inside your own facility rather than a rented datacenter. That is the exact seat ThakiCloud sells. Metis matches a model to the size of the GPU you actually have, and Paxis runs agents on top of it to get real work done. Because the model lives in your rack, the meter does not tick per token. The cloud still has its place. The point is to ask which card is already in your drawer before you rent another.</p>

<hr />

<p><em>An auto-generated comic riffing on this week’s industry news.</em></p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="comics" /><category term="local-ai" /><category term="vram" /><category term="on-prem" /><category term="sovereignty" /><category term="cost" /><category term="thakicloud" /><summary type="html"><![CDATA[We sized the model to the card we already owned. The only thing that starved was the cloud invoice.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://thakicloud.github.io/assets/images/posts/comics/best-local-ai-by-vram-size/strip.png" /><media:content medium="image" url="https://thakicloud.github.io/assets/images/posts/comics/best-local-ai-by-vram-size/strip.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry xml:lang="en"><title type="html">Switch off the OpenAI Realtime API in one line: hugging-voice, an open voice stack you run yourself</title><link href="https://thakicloud.github.io/en/llmops/hugging-voice-open-realtime-voice-self-hosted/" rel="alternate" type="text/html" title="Switch off the OpenAI Realtime API in one line: hugging-voice, an open voice stack you run yourself" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/en/llmops/hugging-voice-open-realtime-voice-self-hosted</id><content type="html" xml:base="https://thakicloud.github.io/en/llmops/hugging-voice-open-realtime-voice-self-hosted/"><![CDATA[<p><img src="/assets/images/hugging-voice-open-realtime-voice-self-hosted-hero.png" alt="An open realtime voice pipeline you run yourself" /></p>

<p>This post is for engineers who wanted to add a voice agent but hesitated at the lock-in and cost of the OpenAI Realtime API, and for infrastructure owners weighing whether conversational voice can be served on their own stack. The short version is that the design of Hugging Face’s demo <a href="https://huggingface.co/spaces/HuggingFaceM4/hugging-voice">hugging-voice</a> and the library beneath it, <a href="https://github.com/huggingface/speech-to-speech">speech-to-speech</a>, is both simple and practical. It opens the entire four-stage realtime voice pipeline as open source, while wrapping the outside in the very same interface as OpenAI Realtime. So if you already have code written against an OpenAI realtime client, you can move onto your own stack by changing a single line: the address the server points to. We only cite performance figures within the range the project has published, and we note up front that these are not numbers we benchmarked ourselves.</p>

<h2 id="overview">Overview</h2>

<p>Over the past year, conversational voice has stopped being a side feature of text chatbots and become a product category of its own. Users speak, and they expect the reply to come back instantly, the way a human conversation flows. The problem is that the commercial path to meeting that expectation has effectively converged on a handful of closed services like the OpenAI Realtime API. Convenient, yes, but voice traffic tends to be billed by the second rather than by the token, data leaves your walls, and both the model and the voices are tied to the provider.</p>

<p>hugging-voice arrives as a counterexample to that trend. The subtitle of the Space says it directly: “An Open Realtime Voice You Can Actually Run Yourself.” The core idea is that the whole round trip, turning the voice arriving at the microphone into text, sending it to a language model, and turning the reply back into speech, is opened as a pipeline whose every component can be swapped. For those of us who serve models in on-prem and sovereign environments, it means there is now a concrete reference implementation for the question “can realtime voice run on our own cluster?”</p>

<h2 id="what-hugging-voice-and-speech-to-speech-are">What hugging-voice and speech-to-speech are</h2>

<p>To settle the terms first: hugging-voice is the demo Space where you can speak to it directly in the browser, and the engine that actually processes the voice inside it is the speech-to-speech library. The library splits a realtime voice agent into four stages: voice activity detection (VAD), speech-to-text (STT), a language model (LLM), and text-to-speech (TTS). Each stage runs in a separate thread and is connected by queues, so the output of one stage streams into the next. A partial transcript appears before the user has finished speaking, and speech synthesis begins on the opening words before the model has finished the sentence, which cuts perceived latency.</p>

<div class="mermaid">
flowchart TB
    A["Microphone input<br />realtime audio stream"] --&gt; B["VAD speech detection<br />Silero VAD v5"]
    B --&gt; C["STT recognition<br />Parakeet TDT · Whisper etc."]
    C --&gt; D["LLM response<br />OpenAI-compatible API · vLLM · llama.cpp"]
    D --&gt; E["TTS synthesis<br />Qwen3-TTS · Kokoro etc."]
    E --&gt; F["Speaker output<br />streaming playback"]
    G["OpenAI Realtime compatible<br />WebSocket server"] -.- B
    G -.- C
    G -.- D
    G -.- E
</div>

<p>In this diagram, the WebSocket server on the right is the project’s real weapon. Simply gluing four stages together is not new. What sets speech-to-speech apart is that it wraps the entire pipeline in a WebSocket endpoint compatible with the OpenAI Realtime protocol. That lets an existing OpenAI realtime client connect to this server as if it were OpenAI itself. It is worth noting that this stack is not an experimental toy: it runs the realtime voice infrastructure of Hugging Face’s Reachy Mini robots in production.</p>

<h2 id="switching-over-in-one-line">Switching over in one line</h2>

<p>This is the part the project itself calls “one-line migration.” The only thing a client that used OpenAI Realtime has to change is the connection address. Below is the Python client example the project documentation offers.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="n">openai</span> <span class="kn">import</span> <span class="n">OpenAI</span>

<span class="n">client</span> <span class="o">=</span> <span class="nc">OpenAI</span><span class="p">(</span>
    <span class="n">base_url</span><span class="o">=</span><span class="sh">"</span><span class="s">http://localhost:8765/v1</span><span class="sh">"</span><span class="p">,</span>
    <span class="n">websocket_base_url</span><span class="o">=</span><span class="sh">"</span><span class="s">ws://localhost:8765/v1</span><span class="sh">"</span><span class="p">,</span>
    <span class="n">api_key</span><span class="o">=</span><span class="sh">"</span><span class="s">not-needed</span><span class="sh">"</span><span class="p">,</span>
<span class="p">)</span>

<span class="k">with</span> <span class="n">client</span><span class="p">.</span><span class="n">realtime</span><span class="p">.</span><span class="nf">connect</span><span class="p">(</span><span class="n">model</span><span class="o">=</span><span class="sh">"</span><span class="s">local</span><span class="sh">"</span><span class="p">)</span> <span class="k">as</span> <span class="n">conn</span><span class="p">:</span>
    <span class="n">conn</span><span class="p">.</span><span class="nf">send</span><span class="p">({</span>
        <span class="sh">"</span><span class="s">type</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">session.update</span><span class="sh">"</span><span class="p">,</span>
        <span class="sh">"</span><span class="s">session</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
            <span class="sh">"</span><span class="s">type</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">realtime</span><span class="sh">"</span><span class="p">,</span>
            <span class="sh">"</span><span class="s">instructions</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">You are a helpful assistant.</span><span class="sh">"</span><span class="p">,</span>
            <span class="sh">"</span><span class="s">audio</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
                <span class="sh">"</span><span class="s">input</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
                    <span class="sh">"</span><span class="s">turn_detection</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
                        <span class="sh">"</span><span class="s">type</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">server_vad</span><span class="sh">"</span><span class="p">,</span>
                        <span class="sh">"</span><span class="s">interrupt_response</span><span class="sh">"</span><span class="p">:</span> <span class="bp">True</span><span class="p">,</span>
                    <span class="p">}</span>
                <span class="p">}</span>
            <span class="p">},</span>
        <span class="p">}</span>
    <span class="p">})</span>

    <span class="k">for</span> <span class="n">event</span> <span class="ow">in</span> <span class="n">conn</span><span class="p">:</span>
        <span class="nf">print</span><span class="p">(</span><span class="n">event</span><span class="p">.</span><span class="nb">type</span><span class="p">)</span>
</code></pre></div></div>

<p>The parts worth noticing are that <code class="language-plaintext highlighter-rouge">base_url</code> and <code class="language-plaintext highlighter-rouge">websocket_base_url</code> point to a local server and that <code class="language-plaintext highlighter-rouge">api_key</code> is essentially not needed. The instructions passed through <code class="language-plaintext highlighter-rouge">session.update</code>, server-side VAD turn detection, and interrupting a response mid-stream all follow the same schema as OpenAI Realtime. In other words, the application code barely changes, and only the backend moves from an external API to your own server. For teams worried about vendor lock-in, this interface compatibility is on its own the biggest practical value.</p>

<h2 id="install-and-run">Install and run</h2>

<p>The path to standing up a server is just as brief. The default install covers the standard realtime path in one shot.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install </span>speech-to-speech
</code></pre></div></div>

<p>The default configuration uses Parakeet TDT for STT, an OpenAI-compatible API for the LLM, and Qwen3-TTS for TTS. If you need a specific backend, install it with extras.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install</span> <span class="s2">"speech-to-speech[kokoro]"</span>
pip <span class="nb">install</span> <span class="s2">"speech-to-speech[faster-whisper]"</span>
</code></pre></div></div>

<p>Running the server looks like this. The command launches an OpenAI Realtime compatible server over a local WebSocket.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">export </span><span class="nv">OPENAI_API_KEY</span><span class="o">=</span>...
speech-to-speech
</code></pre></div></div>

<p>The interesting point here is that the LLM stage can run fully local. Below is an example that stands up a Gemma 4 class model with llama.cpp and has speech-to-speech point at that local endpoint.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>llama-server <span class="nt">-hf</span> ggml-org/gemma-4-E4B-it-GGUF <span class="nt">-np</span> 2 <span class="nt">-c</span> 65536

speech-to-speech <span class="se">\</span>
    <span class="nt">--model_name</span> <span class="s2">"ggml-org/gemma-4-E4B-it-GGUF"</span> <span class="se">\</span>
    <span class="nt">--responses_api_base_url</span> <span class="s2">"http://127.0.0.1:8080/v1"</span> <span class="se">\</span>
    <span class="nt">--responses_api_api_key</span> <span class="s2">""</span>
</code></pre></div></div>

<p>On Apple Silicon Macs you can turn on optimized settings and attach an mlx model, and conversely, if local GPU is short, you can point the LLM backend at Hugging Face Inference Providers.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>speech-to-speech <span class="se">\</span>
    <span class="nt">--local_mac_optimal_settings</span> <span class="se">\</span>
    <span class="nt">--model_name</span> <span class="s2">"mlx-community/Qwen3-4B-Instruct-2507-bf16"</span>
</code></pre></div></div>

<p>Being able to move the same pipeline across the spectrum, from fully local execution to delegated cloud inference, with a handful of flags is a virtue of the design.</p>

<h2 id="modular-swaps-stt-llm-tts-backends">Modular swaps: STT, LLM, TTS backends</h2>

<p>The reason this project reads as a reference architecture rather than a mere demo is that the backend of each stage can be swapped. In summary:</p>

<table>
  <thead>
    <tr>
      <th>Stage</th>
      <th>Default backend</th>
      <th>Alternatives</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>VAD</td>
      <td>Silero VAD v5</td>
      <td>built-in only</td>
    </tr>
    <tr>
      <td>STT</td>
      <td>Parakeet TDT</td>
      <td>Whisper, Faster Whisper, Paraformer</td>
    </tr>
    <tr>
      <td>LLM</td>
      <td>OpenAI-compatible API</td>
      <td>Transformers, mlx-lm, vLLM, llama.cpp</td>
    </tr>
    <tr>
      <td>TTS</td>
      <td>Qwen3-TTS</td>
      <td>Kokoro, Pocket TTS, ChatTTS, MMS</td>
    </tr>
  </tbody>
</table>

<p>There are also four run modes. The default, realtime, is the WebSocket that speaks the OpenAI Realtime protocol; local attaches directly to the microphone and speaker; and websocket and socket exchange raw PCM audio over WebSocket and TCP respectively. Language can be specified or left to auto-detection. Being able to combine STT accuracy, LLM quality and cost, and TTS timbre and latency to fit your own requirements is exactly the degree of freedom a closed API does not give you.</p>

<h2 id="latency-and-a-real-world-case">Latency, and a real-world case</h2>

<p>In a voice agent, latency matters as much as accuracy. People feel a conversation break when the reply is even a few hundred milliseconds late. The <a href="https://huggingface.co/blog/cerebras-gemma4-voice-ai">realtime voice demo</a> that Hugging Face published together with Cerebras targets this latency problem head-on. The configuration uses Nvidia’s Parakeet for STT, Google DeepMind’s Gemma 4 running on Cerebras inference for the language model, and Alibaba’s Qwen3-TTS for TTS. The goal is to drive down the LLM stage’s response time with Cerebras’ ultra-fast inference so that the conversation flows as naturally as talking to a person. That said, within what this article could confirm, no specific millisecond figures were published, so we hold off on a quantitative comparison.</p>

<p>There is production evidence too. This stack drives the Reachy Mini robots mentioned earlier, and Hugging Face states that more than 9,000 robots are already deployed in the field. The reason the project obsesses over latency is that in embedded settings, responsiveness is what makes an interaction “feel alive.” We have separately covered the latency budget of voice agents from a GPU serving angle, so if you are thinking about how to allocate latency across each stage of the pipeline, we suggest reading <a href="/en/llmops/voice-agent-latency-budget-gpu-serving/">Voice agent latency budget and GPU serving</a> alongside this.</p>

<h2 id="implications-for-thakiclouds-products">Implications for ThakiCloud’s products</h2>

<p>This pipeline meshes naturally with both of our products.</p>

<p>Through the ai-platform lens, the LLM stage of speech-to-speech ultimately just consumes an OpenAI-compatible endpoint, so you can drop ThakiCloud’s vLLM serving straight into that slot. The STT and TTS models each become separate GPU-occupying workloads, and our strength is exactly in loading such heterogeneous inference workloads onto one cluster together, queuing GPUs with Kueue and isolating them across tenants. Voice traffic tends to accumulate cost by the second, so the unit-cost competitiveness of self-serving works especially strongly, and this open stack fits on-prem and sovereign requirements where data must not leave the premises. To customers for whom a closed voice API is a burden, we can offer the option of “running it on your own cluster through the same interface.”</p>

<p>Through the Paxis lens, voice is a new input and output channel attached to an agent. Paxis is an Agent-Native Cloud control plane that runs on top of ai-platform and treats Skills, Tools, Policies, and Audit Logs as first-class resources, and the OpenAI Realtime compatible interface of speech-to-speech makes it easy to layer this voice channel onto existing agent orchestration. A flow where a spoken instruction becomes the agent’s input and the result of a skill execution returns as speech can be composed while passing through policy gates and audit logs. Low-cost self-serving (ai-platform) creates the economics of voice agents, and on top of it the Agent-Native control plane (Paxis) handles voice as a channel safely.</p>

<h2 id="closing">Closing</h2>

<p>The message hugging-voice sends is clear. Realtime voice no longer has to lean solely on a few providers’ closed APIs; when you need to, you can keep the interface as is and move only the backend onto your own infrastructure. This design, picking each stage from VAD through TTS and changing just one line in the client, sharply lowers the barrier for teams considering self-serving. If you want to see it for yourself, speak to it directly in the <a href="https://huggingface.co/spaces/HuggingFaceM4/hugging-voice">demo Space</a>, or start with <code class="language-plaintext highlighter-rouge">pip install speech-to-speech</code> from the <a href="https://github.com/huggingface/speech-to-speech">GitHub repository</a>.</p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="llmops" /><category term="speech-to-speech" /><category term="realtime-voice" /><category term="VoiceAI" /><category term="OpenAIRealtime" /><category term="STT" /><category term="TTS" /><category term="LLM-serving" /><category term="LLMOps" /><category term="on-prem" /><category term="self-hosting" /><summary type="html"><![CDATA[Hugging Face's hugging-voice and its engine speech-to-speech wrap a full realtime voice pipeline, from VAD through STT, LLM, and TTS, behind an OpenAI Realtime compatible WebSocket. Change one line in your client's base URL and you move onto your own infrastructure. We took the design apart from a serving perspective.]]></summary></entry><entry xml:lang="en"><title type="html">Hugging Face Wasn’t Breached by a Human, but by an Autonomous AI Agent: When the Dataset Pipeline Became the Attack Surface</title><link href="https://thakicloud.github.io/en/news/huggingface-agentic-ai-breach/" rel="alternate" type="text/html" title="Hugging Face Wasn’t Breached by a Human, but by an Autonomous AI Agent: When the Dataset Pipeline Became the Attack Surface" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/en/news/huggingface-agentic-ai-breach</id><content type="html" xml:base="https://thakicloud.github.io/en/news/huggingface-agentic-ai-breach/"><![CDATA[<p><img src="/assets/images/huggingface-agentic-ai-breach-hero.png" alt="Abstract image of an autonomous agent swarm infiltrating a data pipeline" /></p>

<p>The story that shook timelines over the weekend was not a new model or a new benchmark. It was a notice that Hugging Face, the center of the open AI ecosystem, had been breached. What drew even more attention was who did it. According to the company, no human hacker typed commands through the night. An autonomous AI agent framework drove the attack from start to finish.</p>

<p>A company that sells models getting hit by a model makes for a striking narrative. But the point of this post is not to enjoy the irony. For a company like ThakiCloud that handles models and data on top of customer infrastructure, the real work is to soberly separate exactly where the attack entered and what has been confirmed. And the entry point here was not some flashy zero-day. It was the thing we touch every day: a dataset.</p>

<h2 id="what-happened">What Happened</h2>

<p>Hugging Face disclosed the breach in a blog post on Thursday, July 16, 2026. It came after the company had already confirmed unauthorized access to internal datasets and credentials earlier that week and had contained the intrusion. By the company’s account, the intrusion began in the data-processing pipeline, where the attacker used a single malicious dataset to open two code-execution paths.</p>

<p>That is the confirmed skeleton: an autonomous agent drove it, the entry point was a dataset, and two vulnerabilities led to code execution. The remaining details are emphasized differently across outlets, so confirmed facts and secondary reporting should be read apart.</p>

<h2 id="the-attack-path-the-dataset-pipeline-was-the-attack-surface">The Attack Path: The Dataset Pipeline Was the Attack Surface</h2>

<p>The essence is the entry method. The attacker uploaded a malicious dataset to the Hugging Face Hub. The moment that dataset passed through the processing pipeline, two vulnerabilities fired in sequence. One was a remote-code dataset loader path; the other was a template injection while parsing the dataset configuration. Both ultimately resolved into arbitrary code execution.</p>

<p>The idea that a dataset can run code may sound unfamiliar, but practitioners know the risk well. Many dataset loaders trust and execute loading scripts from remote repositories and render configuration fields as templates. That flexibility, built for convenience, becomes an execution channel the moment it meets input that crosses a trust boundary.</p>

<p>What followed once code execution was secured was a textbook breach chain. The attacker escalated with node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over the weekend. The entry was a single point, but from the moment that point granted execution privileges, the spread propagated automatically.</p>

<pre><code class="language-mermaid">flowchart TB
    A[Attacker: uploads malicious dataset] --&gt; B[Dataset processing pipeline]
    B --&gt; C1["Vulnerability 1&lt;br/&gt;remote-code dataset loader"]
    B --&gt; C2["Vulnerability 2&lt;br/&gt;dataset config template injection"]
    C1 --&gt; D[Arbitrary code execution RCE]
    C2 --&gt; D
    D --&gt; E[Node-level access obtained]
    E --&gt; F[Cloud and cluster credentials harvested]
    F --&gt; G[Lateral movement into internal clusters]
    G --&gt; H["Autonomous agent framework&lt;br/&gt;thousands of actions across a swarm of short-lived sandboxes"]
</code></pre>

<h2 id="the-weight-of-saying-an-autonomous-agent-drove-it">The Weight of Saying an Autonomous Agent Drove It</h2>

<p>The novel part of this incident is not the tooling but the cockpit. Hugging Face described the campaign as “an autonomous agent framework executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.” Instead of a human intervening at each step, the agent handled reconnaissance, execution, and movement in a continuous chain.</p>

<p>The problem this structure poses for defenders is speed and scale. A human attacker has physical limits of fatigue and typing speed, but an agent swarm throws thousands of attempts in parallel and moves to the next one the instant a step fails. Using and discarding short-lived sandboxes erases the anchors for detection, and command-and-control that migrates across public services defeats blocklists.</p>

<p>One interesting side note circulated in secondary reporting. As the response unfolded, when the team tried to hand forensics to commercial frontier models (GPT, Claude), safety guardrails reportedly recognized the exploit payloads and command-and-control artifacts as attacks and refused to cooperate, so the team continued detection and analysis with a GLM 5.2-class model [estimated]. This detail comes from some outlets rather than Hugging Face’s official notice, so it is safer not to read it as settled fact. Regardless of its accuracy, though, the tension itself, where a defender cannot use a tool because of its safety policy, is worth recording as something that may recur.</p>

<h2 id="what-was-safe-and-what-is-still-under-investigation">What Was Safe and What Is Still Under Investigation</h2>

<p>The easier an incident is to exaggerate, the clearer the boundaries must be drawn. Hugging Face said it closed the vulnerable code-execution paths, evicted the attacker, rebuilt the compromised nodes, and revoked and rotated all affected credentials. It added that it found no evidence of tampering with public models, user-facing datasets, or Spaces, and that its software supply chain, including container images and published packages, was verified clean.</p>

<p>The user action was a precautionary recommendation. The company advised users to rotate access tokens and review recent account activity. There is an important distinction here. That recommendation is not a confirmation that user tokens were leaked en masse, but a conservative safety measure given the nature of an incident where internal credentials were harvested. Whether partner or customer data was affected was, as of the disclosure, still under investigation.</p>

<p>In short, what is confirmed is the internal breach and credential theft, the existence of two dataset vulnerabilities, and the swift containment and rotation. What remains open is whether partner and customer data was affected, and the confirmation of some details in secondary reporting (the exact action count, the model-refusal anecdote). Mixing the confirmed with the unconfirmed makes an incident look bigger or smaller than it is.</p>

<h2 id="the-thakicloud-view-treating-dataset-processing-as-a-trust-boundary">The ThakiCloud View: Treating Dataset Processing as a Trust Boundary</h2>

<p>The lesson this incident offers an infrastructure company is clear. A dataset is not a passive file but an active input that can execute code the moment it is processed. So we look at this through two lenses.</p>

<p><strong>Through the ai-platform lens</strong>, ThakiCloud’s ai-platform is a K8s-based multi-tenant AI/ML infrastructure. In such an environment, dataset loading and preprocessing must be treated as input from outside the trust boundary, not inside it. Concretely, this means running dataset-processing jobs in least-privilege isolated containers, blocking network egress by default, and separating node and cloud credentials so workloads cannot touch them directly. That this breach spread from node-level access to credential theft shows again why execution isolation and credential separation must be a default, not an option. This is also why demand for on-prem and sovereign AI is high: the more data and execution stay inside the customer boundary, the smaller the blast radius of such pipeline attacks.</p>

<p><strong>Through the Paxis lens</strong>, this incident overlaps exactly with the threat model that an Agent-Native Cloud is designed for in the first place. Paxis is ThakiCloud’s Agent-Native Cloud, and it treats running skills and tools in isolated sandboxes and passing every action through a policy gate and audit log as first-class principles. That the attacker threw thousands of actions with an autonomous agent swarm proves precisely why a structure that screens agent behavior with policy before execution and records it in an audit log after execution is necessary. To counter an attack pattern that uses and discards short-lived sandboxes, the defender too must isolate each execution, explicitly scope its permissions, and leave a reversible audit trail. Isolated execution plus policy-and-audit is not a luxury of the agent era but a minimum requirement.</p>

<p>The two lenses complement each other. ai-platform narrows the blast radius at the infrastructure layer of dataset processing, while Paxis screens each action at the control layer of agent behavior. In an attack like this one, where entry is a data pipeline and the spread is an autonomous agent, defense at both layers is needed to break the chain.</p>

<h2 id="limits-and-counterpoints">Limits and Counterpoints</h2>

<p>To avoid overconfidence in this post’s conclusions, a few things should be made clear. First, the details of the incident are still being settled. Colorful details like the exact action count, the scope of credential theft, and the commercial-model refusal anecdote lean heavily on secondary reporting and must be distinguished from the confirmed facts of the official notice.</p>

<p>Second, our defensive narrative does not mean complete safety. Isolation and policy-and-audit are design principles that shrink the blast radius, not magic that eliminates the vulnerabilities themselves. Vulnerabilities like remote code execution in a dataset loader or injection in config parsing must continue to be found and patched at the code level, and isolation is the second line of defense that contains the damage when such a vulnerability fires.</p>

<p>Third, overrating autonomous-agent attacks is also risky. The root cause of this breach was not sophisticated AI but two familiar vulnerabilities that let input crossing a trust boundary execute code. The agent was merely the automation that exploited those vulnerabilities faster and wider. So the priority for response still lies in the fundamentals: separating untrusted input from execution privileges, detaching credentials from workloads, and making every execution observable.</p>

<p>Hugging Face’s swift containment and transparent disclosure will stand as a good response example. What remains our homework is simple: treat datasets as code rather than files, and make every agent action a subject of screening and audit.</p>

<h2 id="sources">Sources</h2>

<ul>
  <li><a href="https://huggingface.co/blog/security-incident-july-2026">Security incident disclosure, July 2026 (Hugging Face official blog)</a></li>
  <li><a href="https://www.helpnetsecurity.com/2026/07/20/hugging-face-breached-by-autonomous-ai-agent/">Hugging Face breached by autonomous AI agent (Help Net Security)</a></li>
  <li><a href="https://www.bleepingcomputer.com/news/security/hugging-face-breach-autonomous-ai-agent-system-internal-datasets-credentials/">Hugging Face warns an autonomous AI agent hacked its network (BleepingComputer)</a></li>
  <li><a href="https://thehackernews.com/2026/07/worlds-largest-ai-model-repository.html">World’s Largest AI Model Repository Hugging Face Breached by Autonomous AI Agent (The Hacker News)</a></li>
  <li>Secondary reporting (the exact action count and model-refusal anecdote are cited reporting, not confirmed fact): Cryptobriefing, Undercode Testing</li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="news" /><category term="security" /><category term="huggingface" /><category term="ai-agent" /><category term="supply-chain" /><category term="sandbox" /><category term="dataset-security" /><category term="news" /><category term="thakicloud" /><summary type="html"><![CDATA[In July 2026 Hugging Face disclosed an internal breach driven by an autonomous AI agent. The entry point was a single malicious dataset, and two vulnerabilities in the dataset-processing pipeline led to code execution. We separate what is confirmed from what is still under investigation, and explain why dataset processing must be treated as a trust boundary.]]></summary></entry><entry xml:lang="en"><title type="html">Why Dynamic Batch Tuning Failed to Beat Static Scheduling: An Honest Failure Analysis of a Multi-Tenant vLLM Serving Control Loop</title><link href="https://thakicloud.github.io/en/research/agent-dynamic-batch-tuning-vllm/" rel="alternate" type="text/html" title="Why Dynamic Batch Tuning Failed to Beat Static Scheduling: An Honest Failure Analysis of a Multi-Tenant vLLM Serving Control Loop" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/en/research/agent-dynamic-batch-tuning-vllm</id><content type="html" xml:base="https://thakicloud.github.io/en/research/agent-dynamic-batch-tuning-vllm/"><![CDATA[<p>If you operate a vLLM serving cluster on Kubernetes where multiple tenants share high-end GPUs such as the H200, and you have assumed that Kueue’s static admission control is “good enough,” this post is for you. It is even more worth reading if you have been assuming that “an LLM agent adjusting serving parameters in real time is always better.” This paper shows that assumption is not always true, and it does so with quite concrete reasons.</p>

<h2 id="the-problem-admission-is-static-but-load-is-real-time">The Problem: Admission Is Static, but Load Is Real-Time</h2>

<p>Serving LLM inference for multiple tenants on shared GPUs has become a common operational challenge. A single high-capacity accelerator such as the H200 hosts several tenants at once, each with different traffic burstiness, sequence lengths, and SLOs. On Kubernetes-based platforms, a quota- and queue-based scheduler such as Kueue typically decides which jobs are admitted into the pool, but the serving-time parameters that determine how already-admitted requests actually share the engine’s finite capacity, such as batch size and the maximum number of concurrent sequences, are usually fixed statically at deployment time.</p>

<p>The paper raises the question that naturally follows from this setup. In a multi-tenant H200 GPU pool managed by Kueue, when multiple tenants share a vLLM inference server, can an LLM agent that observes real-time GPU memory and latency telemetry and re-tunes each tenant’s batch size and maximum concurrent sequence count online empirically improve the throughput, p99 latency, and cost Pareto frontier over static Kueue admission control? To answer this question, the authors designed a complete empirical protocol meant to run on a single H200, but they state up front that execution itself was called off because access to the target Kubernetes cluster context was unavailable. As a result, not a single throughput or latency number measured on real hardware appears anywhere in this paper. Instead, the paper delivers three things: a control-loop formalization of the online batch and concurrency tuning problem, a reproducible protocol that can be executed the moment cluster access is restored, and a deterministic queuing simulation that stress-tests the structure of three control policies.</p>

<h2 id="formalizing-the-tuning-problem-as-a-control-loop">Formalizing the Tuning Problem as a Control Loop</h2>

<p>The paper formalizes online per-tenant batch and concurrency tuning as a discrete-time closed-loop control problem with a structure similar to a Markov decision process. The state consists of per-tenant queue depth, GPU memory utilization, recent latency percentiles (p50/p95/p99), and the current concurrency cap. The action is a bounded-width per-tenant increment or decrement to concurrency. The reward function subtracts a penalty for p99 latency that exceeds the SLO and a cost term from throughput, and is designed to jointly optimize throughput, latency, and cost within a safe range. Building on this formalization, the authors also propose how the agent would integrate into an actual deployment. The agent does not replace Kueue’s job admission; it operates as a sidecar inside it, adjusting only how already-admitted jobs divide the engine’s capacity. The key point is that the lever being pulled is not the vLLM engine’s internal <code class="language-plaintext highlighter-rouge">max_num_seqs</code>, but client-side per-tenant admission concurrency. Production vLLM does not expose engine parameters as a knob that can be changed in real time without a restart, so this design reflects the practical constraint that the only point actually adjustable is the number of concurrent requests arriving at the server.</p>

<p><img src="/assets/images/posts/research/agent-dynamic-batch-tuning-vllm/fig-control-loop.png" alt="Control Loop: Agent-Driven Tuning Architecture" />
<em>The control-loop structure in which an LLM agent observes GPU telemetry and per-tenant queue depth, produces a bounded-width concurrency adjustment, and thereby re-tunes the batch configuration. This is a conceptual architecture diagram, not a result measured on hardware.</em></p>

<h2 id="the-warning-the-simulation-raised-naive-dynamic-control-was-worse-than-static">The Warning the Simulation Raised: Naive Dynamic Control Was Worse Than Static</h2>

<p>To fill the gap left by the unexecuted empirical protocol, the authors built a seed-fixed, discrete-time queuing simulation written entirely with Python’s standard library. On a simplified model where two tenants share a finite pool of slots, they ran three policies for 20 seeds of 1,800 seconds each and averaged the results. The static policy fixes each tenant’s concurrency cap at 4. The naive dynamic policy observes GPU memory utilization and 6-second-window p99 latency every 3 seconds and adjusts both tenants together by plus or minus 1. The differentiated agent policy is a surrogate model that mimics more sophisticated agent reasoning by independently reflecting each tenant’s backlog ratio.</p>

<p>The results defied expectations. The static policy was best on both throughput (0.686 req/s) and drop count (an average of 1.9 requests), while the naive dynamic policy’s throughput fell to 0.515 req/s and it dropped an average of 309.5 requests. The differentiated agent surrogate model did better than the naive dynamic policy (throughput 0.601 req/s, 154.7 drops) but still could not catch up to the static policy, and p99 latency was also higher for both dynamic variants than for the static policy (11.75 seconds), at 12.79 and 13.25 seconds respectively.</p>

<p><img src="/assets/images/posts/research/agent-dynamic-batch-tuning-vllm/fig-throughput-dropped.png" alt="Throughput vs. Dropped Requests by Policy" />
<em>The static policy recorded the highest throughput and fewest drops, while both dynamic variants underperformed in the simulation. These are results from a deterministic queuing simulation averaged over 20 seeds, not values measured on an actual GPU.</em></p>

<p><img src="/assets/images/posts/research/agent-dynamic-batch-tuning-vllm/fig-p99-latency.png" alt="p99 Latency by Policy" />
<em>The static policy had the lowest p99 latency, and neither dynamic controller reduced tail latency in this simulation regime. These are simulation results averaged over 20 seeds, not figures measured on hardware.</em></p>

<p>The authors verify that this result is the model’s actual dynamics rather than a simple bug, and trace the cause through four steps. First, because service time follows an exponential distribution, the tail is inherently heavy. An exponential distribution with a mean of 2.5 seconds forms a p99 around 11.5 seconds even with zero queuing at all, which is barely different from the 11.75 seconds the static policy recorded. In other words, most of the observed tail latency comes not from congestion but from the variance of the service time itself. Second, the throttle threshold is set at 6 seconds, twice the 3-second window, which is far below this intrinsic p99 of 11.5 seconds. The short-window p99, estimated from only a handful of completed requests, is noisy and biased toward this intrinsic tail, so it frequently crosses the threshold even under genuinely light load. Third, because the resulting false-positive throttles lower both tenants together, even a non-congested tenant gets restricted along with the congested one, and under a 3-second hard timeout, every unnecessary throttle converts directly into a drop. Fourth, while the scaling rule does allow recovery, the throttle-drop cost is asymmetrically larger than the scaling gain, so net throughput falls.</p>

<h2 id="design-requirements-drawn-from-the-failure">Design Requirements Drawn from the Failure</h2>

<p>What this diagnosis tells us is not that dynamic tuning itself is useless. It is that when a noisy, tail-biased signal, throttling that couples tenants together, and a hard timeout combine, naive fixed-threshold control can actually turn out worse than static control. From here, the authors derive four design requirements that an effective agent controller must satisfy. The congestion signal must be robustly estimated from longer telemetry history and confidence estimates rather than raw short-window percentiles. Decisions must be differentiated per tenant rather than driven by a global signal that lumps tenants together. The asymmetric cost between throttling and scaling must be explicitly accounted for, and additional context such as burst-pattern recognition or per-tenant SLA metadata, which cannot be expressed with a fixed numeric threshold, must also be leveraged. The fact that the differentiated agent surrogate model actually recovered roughly half the loss relative to the naive controller suggests this direction may well be valid. However, the authors are clear that because both the coupling and the signal type were changed at the same time, this experiment alone cannot isolate which of the two contributed to the recovery.</p>

<h2 id="what-this-leaves-for-the-company-society-and-science">What This Leaves for the Company, Society, and Science</h2>

<p>For ThakiCloud, what this research leaves us right now is not a verified cost saving. It is a reproducible empirical protocol directly applicable to our multi-tenant H200 cluster, and a concrete warning that a naive approach could actually make things worse. The harness itself is already complete, so that once cluster access is restored, all that needs to be substituted into this protocol is an actual LLM agent decision function. More broadly, as methodologies for reducing resource waste in shared GPU pools are refined, a practical path opens up for smaller organizations to also meet SLAs on shared clusters. This benefits the overall energy and cost efficiency of inference infrastructure. Scientifically, the contribution is that it explicitly formalizes the control loop in which an LLM agent observes real-time system telemetry and re-tunes serving hyperparameters online, and quantitatively diagnoses its failure modes through a reproducible simulation. This paper is also part of the research lineage ThakiCloud has continued on the same Kueue-and-GPU foundation. Compared with ABJ-Gate, which schedules the judge model’s sampling budget, Attested Confidential Sovereign Inference, which handles TEE remote attestation at admission time, and “Escalate or Act?”, which handles binary escalate-or-act decisions for rare discrete accidents, this paper is distinguished from them in that it tackles a qualitatively different problem: real-time, multi-objective control that is continuously re-tuned on a second-by-second cadence.</p>

<h2 id="limitations">Limitations</h2>

<p>The limitations this paper itself discloses are clear. The most important is that there is no real-hardware validation of any kind: the empirical protocol was never executed because access to the target cluster context was unavailable, and not a single number in the paper was measured on a GPU. The simulation simplifies tokens and the KV cache down to a single scalar slot count and does not model prefill/decode or memory-pressure dynamics, and even the memory-utilization signal is merely a proxy for slot occupancy rather than an actual device-memory measurement. The tenant types tested are also limited to two synthetic traffic patterns, so the failure modes discovered may not generalize to other traffic mixes. The agent policy inside the simulation is likewise a hand-written surrogate model rather than an actual LLM, and because both the coupling and the signal type were changed at once, there is as yet no verified evidence that an actual LLM agent would satisfy the design requirements derived in Section 4. The authors also state plainly that they reported only the mean and standard deviation across 20 seeds and did not run any test of statistical significance. For these reasons, this paper positions itself not as a validated result but as a formalization, a protocol, and a cautionary simulation.</p>

<p>You can find the paper’s detail page here: <a href="https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-21-agent-dynamic-batch-tuning-vllm">https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-21-agent-dynamic-batch-tuning-vllm</a></p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="research" /><category term="vllm" /><category term="kueue" /><category term="gpu-scheduling" /><category term="multi-tenant-serving" /><category term="llm-agents" /><category term="dynamic-batching" /><category term="inference-cost-optimization" /><category term="h200" /><category term="queuing-simulation" /><category term="control-loop" /><summary type="html"><![CDATA[Does an agent that observes GPU telemetry and adjusts batch size and concurrency in real time outperform static Kueue admission control? The simulation said no, and traced exactly why.]]></summary></entry><entry xml:lang="ko"><title type="html">OpenAI 실시간 음성 API를 한 줄로 갈아타기: 직접 돌리는 오픈 음성 스택 hugging-voice</title><link href="https://thakicloud.github.io/ko/llmops/hugging-voice-open-realtime-voice-self-hosted/" rel="alternate" type="text/html" title="OpenAI 실시간 음성 API를 한 줄로 갈아타기: 직접 돌리는 오픈 음성 스택 hugging-voice" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/ko/llmops/hugging-voice-open-realtime-voice-self-hosted</id><content type="html" xml:base="https://thakicloud.github.io/ko/llmops/hugging-voice-open-realtime-voice-self-hosted/"><![CDATA[<p><img src="/assets/images/hugging-voice-open-realtime-voice-self-hosted-hero.png" alt="직접 돌리는 오픈 실시간 음성 파이프라인" /></p>

<p>이 글은 음성 에이전트를 붙이려다 OpenAI Realtime API의 종속과 비용 앞에서 멈칫한 엔지니어, 그리고 대화형 음성 기능을 자체 인프라에서 서빙할 수 있을지 저울질하는 인프라 담당자를 위해 썼습니다. 결론부터 말하면, 허깅페이스가 공개한 데모 <a href="https://huggingface.co/spaces/HuggingFaceM4/hugging-voice">hugging-voice</a>와 그 밑에서 돌아가는 라이브러리 <a href="https://github.com/huggingface/speech-to-speech">speech-to-speech</a>의 설계는 단순하면서도 실용적입니다. 실시간 음성을 위한 네 단계 파이프라인을 그대로 오픈소스로 열어 두되, 바깥쪽은 OpenAI Realtime과 똑같은 인터페이스로 감쌌습니다. 그래서 이미 OpenAI 실시간 클라이언트로 짜 둔 코드가 있다면, 서버를 가리키는 주소 한 줄만 바꿔 자체 스택으로 옮겨올 수 있습니다. 다만 성능에 관한 수치는 프로젝트가 공개한 범위에서만 인용하고, 저희가 직접 벤치마크한 값이 아님을 먼저 분명히 해 둡니다.</p>

<h2 id="개요">개요</h2>

<p>지난 1년 사이 대화형 음성은 텍스트 챗봇의 곁가지가 아니라 독립된 제품군이 됐습니다. 사용자는 말을 걸고, 답이 사람과 대화하듯 즉시 돌아오기를 기대합니다. 문제는 이 기대치를 맞추는 상용 경로가 사실상 OpenAI Realtime API 같은 소수의 폐쇄형 서비스로 수렴해 왔다는 점입니다. 편하긴 하지만 음성 트래픽은 토큰이 아니라 초 단위로 과금되기 쉽고, 데이터가 외부로 나가며, 모델과 목소리 선택지가 공급자에 묶입니다.</p>

<p>hugging-voice는 이 흐름에 대한 반례로 등장했습니다. 스페이스의 부제부터가 “직접 돌릴 수 있는 오픈 실시간 음성(An Open Realtime Voice You Can Actually Run Yourself)”입니다. 마이크로 들어온 목소리를 텍스트로 바꾸고, 언어 모델에 보내고, 답을 다시 음성으로 되돌려 주는 이 왕복 전체를, 구성 요소 하나하나를 교체할 수 있는 오픈 파이프라인으로 열어 둔 것이 핵심입니다. 저희처럼 온프렘과 소버린 환경에서 모델을 서빙하는 입장에서는, “실시간 음성을 자체 클러스터에서 돌릴 수 있는가”라는 질문에 대한 구체적인 참조 구현이 생긴 셈입니다.</p>

<h2 id="hugging-voice와-speech-to-speech는-무엇인가">hugging-voice와 speech-to-speech는 무엇인가</h2>

<p>용어부터 정리하면, hugging-voice는 웹에서 바로 말을 걸어 볼 수 있는 데모 스페이스이고, 그 안에서 실제로 음성을 처리하는 엔진이 speech-to-speech 라이브러리입니다. 라이브러리는 실시간 음성 에이전트를 네 단계로 나눕니다. 음성 구간 감지(VAD), 음성 인식(STT), 언어 모델(LLM), 음성 합성(TTS)입니다. 각 단계는 별도 스레드에서 돌고 큐로 연결되어, 앞 단계의 출력이 스트리밍으로 다음 단계에 흘러 들어갑니다. 사용자가 말을 마치기도 전에 부분 전사가 나오고, 모델이 문장을 완성하기 전에 앞부분부터 음성 합성이 시작되는 구조라 체감 지연을 줄일 수 있습니다.</p>

<div class="mermaid">
flowchart TB
    A["마이크 입력<br />실시간 오디오 스트림"] --&gt; B["VAD 발화 구간 감지<br />Silero VAD v5"]
    B --&gt; C["STT 음성 인식<br />Parakeet TDT · Whisper 등"]
    C --&gt; D["LLM 응답 생성<br />OpenAI 호환 API · vLLM · llama.cpp"]
    D --&gt; E["TTS 음성 합성<br />Qwen3-TTS · Kokoro 등"]
    E --&gt; F["스피커 출력<br />스트리밍 재생"]
    G["OpenAI Realtime 호환<br />WebSocket 서버"] -.- B
    G -.- C
    G -.- D
    G -.- E
</div>

<p>이 그림에서 오른쪽의 WebSocket 서버가 이 프로젝트의 진짜 무기입니다. 네 단계를 잘 붙였다는 것만으로는 새롭지 않습니다. speech-to-speech가 다른 점은, 이 파이프라인 전체를 OpenAI Realtime 프로토콜과 호환되는 WebSocket 엔드포인트로 감쌌다는 데 있습니다. 덕분에 기존 OpenAI 실시간 클라이언트가 이 서버를 마치 OpenAI인 것처럼 붙어서 쓸 수 있습니다. 참고로 이 스택은 실험용 장난감이 아니라, 허깅페이스가 판매하는 Reachy Mini 로봇의 실시간 음성 인프라를 실제로 구동하고 있습니다.</p>

<h2 id="한-줄로-갈아타기">한 줄로 갈아타기</h2>

<p>프로젝트가 스스로 “한 줄 이전(one-line migration)”이라고 부르는 대목이 여기입니다. OpenAI Realtime을 쓰던 클라이언트에서 바꿔야 하는 것은 접속 주소뿐입니다. 아래는 프로젝트 문서가 제시하는 파이썬 클라이언트 예시입니다.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">from</span> <span class="n">openai</span> <span class="kn">import</span> <span class="n">OpenAI</span>

<span class="n">client</span> <span class="o">=</span> <span class="nc">OpenAI</span><span class="p">(</span>
    <span class="n">base_url</span><span class="o">=</span><span class="sh">"</span><span class="s">http://localhost:8765/v1</span><span class="sh">"</span><span class="p">,</span>
    <span class="n">websocket_base_url</span><span class="o">=</span><span class="sh">"</span><span class="s">ws://localhost:8765/v1</span><span class="sh">"</span><span class="p">,</span>
    <span class="n">api_key</span><span class="o">=</span><span class="sh">"</span><span class="s">not-needed</span><span class="sh">"</span><span class="p">,</span>
<span class="p">)</span>

<span class="k">with</span> <span class="n">client</span><span class="p">.</span><span class="n">realtime</span><span class="p">.</span><span class="nf">connect</span><span class="p">(</span><span class="n">model</span><span class="o">=</span><span class="sh">"</span><span class="s">local</span><span class="sh">"</span><span class="p">)</span> <span class="k">as</span> <span class="n">conn</span><span class="p">:</span>
    <span class="n">conn</span><span class="p">.</span><span class="nf">send</span><span class="p">({</span>
        <span class="sh">"</span><span class="s">type</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">session.update</span><span class="sh">"</span><span class="p">,</span>
        <span class="sh">"</span><span class="s">session</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
            <span class="sh">"</span><span class="s">type</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">realtime</span><span class="sh">"</span><span class="p">,</span>
            <span class="sh">"</span><span class="s">instructions</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">You are a helpful assistant.</span><span class="sh">"</span><span class="p">,</span>
            <span class="sh">"</span><span class="s">audio</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
                <span class="sh">"</span><span class="s">input</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
                    <span class="sh">"</span><span class="s">turn_detection</span><span class="sh">"</span><span class="p">:</span> <span class="p">{</span>
                        <span class="sh">"</span><span class="s">type</span><span class="sh">"</span><span class="p">:</span> <span class="sh">"</span><span class="s">server_vad</span><span class="sh">"</span><span class="p">,</span>
                        <span class="sh">"</span><span class="s">interrupt_response</span><span class="sh">"</span><span class="p">:</span> <span class="bp">True</span><span class="p">,</span>
                    <span class="p">}</span>
                <span class="p">}</span>
            <span class="p">},</span>
        <span class="p">}</span>
    <span class="p">})</span>

    <span class="k">for</span> <span class="n">event</span> <span class="ow">in</span> <span class="n">conn</span><span class="p">:</span>
        <span class="nf">print</span><span class="p">(</span><span class="n">event</span><span class="p">.</span><span class="nb">type</span><span class="p">)</span>
</code></pre></div></div>

<p>눈여겨볼 부분은 <code class="language-plaintext highlighter-rouge">base_url</code>과 <code class="language-plaintext highlighter-rouge">websocket_base_url</code>이 로컬 서버를 가리키고, <code class="language-plaintext highlighter-rouge">api_key</code>가 사실상 필요 없다는 점입니다. <code class="language-plaintext highlighter-rouge">session.update</code>로 넘기는 지시문, 서버 측 VAD 기반 턴 감지, 응답 중 끼어들기 같은 옵션은 OpenAI Realtime과 같은 스키마를 그대로 따릅니다. 즉 애플리케이션 코드는 거의 손대지 않고, 백엔드만 외부 API에서 자체 서버로 옮겨오는 구조입니다. 벤더 종속을 걱정하는 팀에게는 이 인터페이스 호환성 자체가 가장 큰 실용적 가치입니다.</p>

<h2 id="설치와-실행">설치와 실행</h2>

<p>서버를 띄우는 경로도 간결합니다. 기본 설치는 표준 실시간 경로를 한 번에 덮습니다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install </span>speech-to-speech
</code></pre></div></div>

<p>기본 구성은 STT로 Parakeet TDT, LLM으로 OpenAI 호환 API, TTS로 Qwen3-TTS를 씁니다. 특정 백엔드가 필요하면 추가 익스트라로 설치합니다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install</span> <span class="s2">"speech-to-speech[kokoro]"</span>
pip <span class="nb">install</span> <span class="s2">"speech-to-speech[faster-whisper]"</span>
</code></pre></div></div>

<p>서버 실행은 다음과 같습니다. 이 명령은 OpenAI Realtime 호환 서버를 로컬 WebSocket으로 띄웁니다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">export </span><span class="nv">OPENAI_API_KEY</span><span class="o">=</span>...
speech-to-speech
</code></pre></div></div>

<p>여기서 흥미로운 지점은, LLM 단계를 완전히 로컬로 돌릴 수 있다는 것입니다. 아래는 llama.cpp로 Gemma 4 계열 모델을 띄우고, speech-to-speech가 그 로컬 엔드포인트를 바라보게 하는 예시입니다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>llama-server <span class="nt">-hf</span> ggml-org/gemma-4-E4B-it-GGUF <span class="nt">-np</span> 2 <span class="nt">-c</span> 65536

speech-to-speech <span class="se">\</span>
    <span class="nt">--model_name</span> <span class="s2">"ggml-org/gemma-4-E4B-it-GGUF"</span> <span class="se">\</span>
    <span class="nt">--responses_api_base_url</span> <span class="s2">"http://127.0.0.1:8080/v1"</span> <span class="se">\</span>
    <span class="nt">--responses_api_api_key</span> <span class="s2">""</span>
</code></pre></div></div>

<p>애플 실리콘 맥이라면 최적화 옵션을 켜고 mlx 모델을 붙일 수 있고, 반대로 로컬 GPU가 부족하면 허깅페이스 인퍼런스 프로바이더를 LLM 백엔드로 지정할 수도 있습니다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>speech-to-speech <span class="se">\</span>
    <span class="nt">--local_mac_optimal_settings</span> <span class="se">\</span>
    <span class="nt">--model_name</span> <span class="s2">"mlx-community/Qwen3-4B-Instruct-2507-bf16"</span>
</code></pre></div></div>

<p>같은 파이프라인을 로컬 완전 실행부터 클라우드 추론 위임까지, 플래그 몇 개로 오갈 수 있다는 점이 설계의 미덕입니다.</p>

<h2 id="모듈-교체-stt-llm-tts-백엔드">모듈 교체: STT, LLM, TTS 백엔드</h2>

<p>이 프로젝트가 단순한 데모를 넘어 참조 아키텍처로 읽히는 이유는 각 단계의 백엔드를 갈아 끼울 수 있기 때문입니다. 정리하면 다음과 같습니다.</p>

<table>
  <thead>
    <tr>
      <th>단계</th>
      <th>기본 백엔드</th>
      <th>교체 가능한 대안</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>VAD</td>
      <td>Silero VAD v5</td>
      <td>내장 전용</td>
    </tr>
    <tr>
      <td>STT</td>
      <td>Parakeet TDT</td>
      <td>Whisper, Faster Whisper, Paraformer</td>
    </tr>
    <tr>
      <td>LLM</td>
      <td>OpenAI 호환 API</td>
      <td>Transformers, mlx-lm, vLLM, llama.cpp</td>
    </tr>
    <tr>
      <td>TTS</td>
      <td>Qwen3-TTS</td>
      <td>Kokoro, Pocket TTS, ChatTTS, MMS</td>
    </tr>
  </tbody>
</table>

<p>실행 모드도 네 가지입니다. 기본값인 realtime은 OpenAI Realtime 프로토콜의 WebSocket이고, local은 마이크와 스피커에 직접 붙는 모드, websocket과 socket은 각각 WebSocket과 TCP로 원시 PCM 오디오를 주고받는 저수준 모드입니다. 언어도 지정하거나 자동 감지에 맡길 수 있습니다. 이렇게 STT 정확도, LLM 품질과 비용, TTS 음색과 지연을 각자의 요구에 맞춰 조합할 수 있다는 점이, 폐쇄형 API가 주지 못하는 자유도입니다.</p>

<h2 id="지연시간-그리고-실제-사례">지연시간, 그리고 실제 사례</h2>

<p>음성 에이전트에서 지연은 정확도만큼이나 중요한 변수입니다. 사람은 응답이 몇백 밀리초만 늦어도 대화가 끊긴다고 느낍니다. 허깅페이스가 세레브라스(Cerebras)와 함께 공개한 <a href="https://huggingface.co/blog/cerebras-gemma4-voice-ai">실시간 음성 데모</a>는 이 지연 문제를 정면으로 겨냥합니다. 구성은 STT에 엔비디아 Parakeet, 언어 모델에 세레브라스 추론 위에서 도는 구글 딥마인드 Gemma 4, TTS에 알리바바 Qwen3-TTS를 쓰는 조합입니다. 세레브라스의 초고속 추론으로 LLM 단계의 응답 시간을 끌어내려, 대화가 사람과의 대화만큼 자연스럽게 흐르도록 만드는 것이 목표입니다. 다만 이 글이 확인한 범위에서 구체적인 밀리초 단위 수치는 공개되지 않았으므로, 정량 비교는 유보합니다.</p>

<p>프로덕션 근거도 있습니다. 이 스택은 앞서 언급한 Reachy Mini 로봇을 구동하며, 허깅페이스는 이미 9,000대 이상의 로봇이 현장에 배포됐다고 밝히고 있습니다. 임베디드 환경에서 반응 속도가 곧 “상호작용이 살아 있다는 느낌”을 좌우한다는 것이, 이 프로젝트가 지연에 집착하는 이유입니다. 저희 블로그에서 음성 에이전트의 지연 예산을 GPU 서빙 관점에서 따로 다룬 적이 있으니, 파이프라인 각 단계의 지연을 어떻게 배분할지 고민한다면 <a href="/ko/llmops/voice-agent-latency-budget-gpu-serving/">음성 에이전트 지연 예산과 GPU 서빙</a> 글을 함께 보시길 권합니다.</p>

<h2 id="thakicloud-제품-적용-시사점">ThakiCloud 제품 적용 시사점</h2>

<p>이 파이프라인은 저희 두 제품 모두와 자연스럽게 맞물립니다.</p>

<p>ai-platform 렌즈에서 보면, speech-to-speech의 LLM 단계는 결국 OpenAI 호환 엔드포인트를 소비할 뿐이므로, 그 자리에 ThakiCloud의 vLLM 서빙을 그대로 꽂을 수 있습니다. STT와 TTS 모델은 각각 GPU를 점유하는 별도 워크로드가 되는데, 저희는 Kueue로 GPU를 큐잉하고 멀티테넌트로 격리해 이런 이종 추론 워크로드를 한 클러스터에 함께 태우는 데 강점이 있습니다. 음성 트래픽은 초 단위로 과금이 불어나기 쉬워 자체 서빙의 단가 경쟁력이 특히 크게 작동하고, 데이터가 외부로 나가지 않아야 하는 온프렘과 소버린 요구에도 이 오픈 스택은 그대로 부합합니다. 폐쇄형 음성 API가 부담스러운 고객에게, “같은 인터페이스로 자체 클러스터에서 돌린다”는 선택지를 제시할 수 있는 것입니다.</p>

<p>Paxis 렌즈에서 보면, 음성은 에이전트에 붙는 새로운 입출력 채널입니다. Paxis는 ai-platform 위에서 도는 Agent-Native Cloud 제어 평면으로 Skills, Tools, Policies, Audit Logs를 일급 리소스로 다루는데, speech-to-speech의 OpenAI Realtime 호환 인터페이스는 이 음성 채널을 기존 에이전트 오케스트레이션에 얹기 쉽게 만듭니다. 사용자가 말로 지시하면 그 발화가 곧 에이전트의 입력이 되고, 스킬 실행 결과가 음성으로 되돌아오는 흐름을, 정책 게이트와 감사 로그를 통과시키면서 구성할 수 있습니다. 저비용 자체 서빙(ai-platform)이 음성 에이전트의 경제성을 만들고, 그 위에서 Agent-Native 제어 평면(Paxis)이 음성이라는 채널을 안전하게 다루는 그림입니다.</p>

<h2 id="마무리">마무리</h2>

<p>hugging-voice가 던지는 메시지는 명확합니다. 실시간 음성은 더 이상 소수 공급자의 폐쇄형 API에만 기댈 필요가 없고, 필요하면 인터페이스는 그대로 둔 채 백엔드만 자체 인프라로 옮겨올 수 있다는 것입니다. VAD에서 TTS까지 각 단계를 골라 끼우고, 클라이언트는 주소 한 줄만 바꾸는 이 설계는, 자체 서빙을 검토하는 팀에게 진입 비용을 크게 낮춰 줍니다. 직접 확인하고 싶다면 <a href="https://huggingface.co/spaces/HuggingFaceM4/hugging-voice">데모 스페이스</a>에서 바로 말을 걸어 보거나, <a href="https://github.com/huggingface/speech-to-speech">GitHub 저장소</a>에서 <code class="language-plaintext highlighter-rouge">pip install speech-to-speech</code>로 시작하시면 됩니다.</p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="llmops" /><category term="speech-to-speech" /><category term="실시간음성" /><category term="VoiceAI" /><category term="OpenAIRealtime" /><category term="STT" /><category term="TTS" /><category term="LLM서빙" /><category term="LLMOps" /><category term="온프렘" /><category term="self-hosting" /><summary type="html"><![CDATA[허깅페이스가 공개한 hugging-voice와 그 엔진인 speech-to-speech는 VAD에서 STT, LLM, TTS까지 이어지는 실시간 음성 파이프라인을 OpenAI Realtime 호환 WebSocket으로 감쌌습니다. 클라이언트의 base URL 한 줄만 바꾸면 자체 인프라로 옮겨올 수 있다는 이 접근을, 서빙 관점에서 뜯어봤습니다.]]></summary></entry><entry xml:lang="ko"><title type="html">허깅페이스를 뚫은 건 사람이 아니라 자율 AI 에이전트였습니다: 데이터셋 파이프라인이 공격면이 된 사건</title><link href="https://thakicloud.github.io/ko/news/huggingface-agentic-ai-breach/" rel="alternate" type="text/html" title="허깅페이스를 뚫은 건 사람이 아니라 자율 AI 에이전트였습니다: 데이터셋 파이프라인이 공격면이 된 사건" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/ko/news/huggingface-agentic-ai-breach</id><content type="html" xml:base="https://thakicloud.github.io/ko/news/huggingface-agentic-ai-breach/"><![CDATA[<p><img src="/assets/images/huggingface-agentic-ai-breach-hero.png" alt="자율 에이전트 스웜이 데이터 파이프라인을 파고드는 추상 이미지" /></p>

<p>지난 주말 타임라인을 흔든 소식은 새 모델도, 새 벤치마크도 아니었습니다. 오픈 AI 생태계의 중심인 허깅페이스가 뚫렸다는 공지였습니다. 더 눈길을 끈 것은 침해의 주체였습니다. 사람 해커가 밤새 손으로 명령을 친 것이 아니라, 자율 AI 에이전트 프레임워크가 공격을 처음부터 끝까지 몰고 갔다고 회사가 밝혔기 때문입니다.</p>

<p>모델을 파는 회사가 모델에게 당했다는 사건의 구도는 자극적입니다. 그러나 이 글의 목적은 그 아이러니를 소비하는 것이 아닙니다. ThakiCloud처럼 고객 인프라 위에서 모델과 데이터를 다루는 입장에서는, 공격의 진입로가 정확히 어디였고 무엇이 확인됐는지를 냉정하게 구분하는 일이 곧 실무입니다. 이번 사건의 진입로는 화려한 제로데이가 아니라 우리가 매일 만지는 것, 바로 데이터셋이었습니다.</p>

<h2 id="무슨-일이-있었나">무슨 일이 있었나</h2>

<p>허깅페이스는 2026년 7월 16일 목요일 블로그를 통해 침해 사실을 공개했습니다. 그 주 초에 내부 데이터셋과 자격증명에 대한 비인가 접근을 확인한 뒤 대응을 마무리하고 나서 낸 발표였습니다. 회사의 설명에 따르면 침입은 데이터 처리 파이프라인에서 시작됐고, 공격자는 악성 데이터셋 하나를 이용해 두 개의 코드 실행 경로를 열었습니다.</p>

<p>여기까지가 발표 주체가 명시적으로 내놓은 뼈대입니다. 자율 에이전트가 몰았다는 점, 진입로가 데이터셋이었다는 점, 그리고 두 개의 취약점이 코드 실행으로 이어졌다는 점입니다. 나머지 세부는 보도마다 강조점이 조금씩 다르므로, 확인된 사실과 이차 보도를 구분해 읽어야 합니다.</p>

<h2 id="공격-경로-데이터셋-파이프라인이-공격면이었다">공격 경로: 데이터셋 파이프라인이 공격면이었다</h2>

<p>핵심은 진입 방식입니다. 공격자는 허깅페이스 허브에 악성 데이터셋을 올렸습니다. 그 데이터셋이 처리 파이프라인을 통과하는 순간, 두 개의 취약점이 연달아 터졌습니다. 하나는 원격 코드 데이터셋 로더 경로였고, 다른 하나는 데이터셋 설정을 파싱하는 과정의 템플릿 인젝션이었습니다. 둘 다 최종적으로는 임의 코드 실행으로 귀결됐습니다.</p>

<p>데이터셋이 코드를 실행시킬 수 있다는 사실이 낯설게 들릴 수 있습니다. 그러나 실무자라면 익숙한 위험입니다. 많은 데이터셋 로더가 원격 저장소의 로딩 스크립트를 신뢰하고 실행하며, 설정 파일의 필드를 템플릿으로 렌더링합니다. 편의를 위해 만든 이 유연성이, 신뢰 경계를 넘는 입력을 만나면 그대로 실행 통로가 됩니다.</p>

<p>코드 실행을 확보한 다음의 전개는 전형적인 침해 체인이었습니다. 공격자는 노드 레벨 접근으로 권한을 끌어올렸고, 클라우드와 클러스터 자격증명을 수집했으며, 주말 동안 여러 내부 클러스터로 측면 이동했습니다. 진입은 한 지점이었지만, 그 지점이 실행 권한을 주는 순간부터 확산은 자동으로 번졌습니다.</p>

<pre><code class="language-mermaid">flowchart TB
    A[공격자: 악성 데이터셋 업로드] --&gt; B[데이터셋 처리 파이프라인]
    B --&gt; C1["취약점 1&lt;br/&gt;remote-code dataset loader"]
    B --&gt; C2["취약점 2&lt;br/&gt;dataset config template injection"]
    C1 --&gt; D[임의 코드 실행 RCE]
    C2 --&gt; D
    D --&gt; E[노드 레벨 접근 획득]
    E --&gt; F[클라우드·클러스터 자격증명 탈취]
    F --&gt; G[내부 클러스터로 측면 이동]
    G --&gt; H["자율 에이전트 프레임워크&lt;br/&gt;단기 샌드박스 스웜에서 수천 개 액션 수행"]
</code></pre>

<h2 id="자율-에이전트가-몰았다는-말의-무게">자율 에이전트가 몰았다는 말의 무게</h2>

<p>이번 사건에서 가장 새로운 대목은 도구가 아니라 조종석입니다. 허깅페이스는 이 캠페인을 “자율 에이전트 프레임워크가 단기 샌드박스 스웜에서 수천 개의 개별 액션을 수행했고, 명령·제어 채널은 공개 서비스 위에 스스로 이전해 가며 자리를 잡았다”고 설명했습니다. 사람이 단계마다 개입하는 대신, 에이전트가 정찰과 실행과 이동을 이어서 처리했다는 뜻입니다.</p>

<p>이 구조가 방어자에게 던지는 문제는 속도와 규모입니다. 사람 공격자라면 피로와 타이핑 속도라는 물리적 한계가 있지만, 에이전트 스웜은 병렬로 수천 개의 시도를 던지고 실패하면 즉시 다음으로 넘어갑니다. 단기 샌드박스를 쓰고 버리는 방식은 탐지의 앵커를 지우며, 공개 서비스 위에서 이전하는 명령·제어는 차단 목록을 무력화합니다.</p>

<p>한 가지 흥미로운 곁가지가 이차 보도에서 돌았습니다. 대응 과정에서 상용 프런티어 모델(GPT, Claude)에 포렌식을 맡기려 하자 안전 가드레일이 익스플로잇 페이로드와 명령·제어 아티팩트를 공격으로 인식해 협조를 거부했고, 결국 GLM 5.2 계열 모델로 탐지와 분석을 이어갔다는 내용입니다 [추정]. 이 대목은 허깅페이스 공식 공지가 아니라 일부 매체 보도에 근거하므로 확정 사실로 읽지 않는 편이 안전합니다. 다만 사실 여부와 별개로, 보안 사고를 방어하는 쪽이 안전 정책 때문에 도구를 못 쓰는 상황 자체가 앞으로 반복될 수 있는 긴장이라는 점은 기록해 둘 만합니다.</p>

<h2 id="무엇이-안전했고-무엇이-아직-조사-중인가">무엇이 안전했고 무엇이 아직 조사 중인가</h2>

<p>과장하기 쉬운 사건일수록 경계선을 분명히 그어야 합니다. 허깅페이스는 취약한 코드 실행 경로를 닫고, 공격자를 축출하고, 침해된 노드를 재구축했으며, 영향을 받은 자격증명을 전부 무효화하고 교체했다고 밝혔습니다. 또한 공개된 모델과 사용자용 데이터셋, Spaces에 변조 흔적은 발견되지 않았고, 컨테이너 이미지와 배포 패키지를 포함한 소프트웨어 공급망은 깨끗한 것으로 검증됐다고 덧붙였습니다.</p>

<p>사용자 조치는 예방 차원의 권고였습니다. 회사는 사용자에게 접근 토큰을 교체하고 최근 계정 활동을 검토하라고 안내했습니다. 여기서 중요한 구분이 있습니다. 이 권고는 사용자 토큰이 대량으로 유출됐다는 확인이 아니라, 내부 자격증명이 탈취된 사고의 성격상 취하는 보수적 안전 조치라는 점입니다. 파트너나 고객 데이터가 영향을 받았는지는 발표 시점 기준으로 여전히 조사 중이라고 회사는 밝혔습니다.</p>

<p>정리하면 확인된 것은 내부 침해와 자격증명 탈취, 두 데이터셋 취약점의 존재, 그리고 신속한 봉쇄와 교체입니다. 아직 열려 있는 것은 파트너·고객 데이터 영향 여부와, 이차 보도에 담긴 일부 세부(정확한 액션 수, 모델 협조 거부 일화)의 확정입니다. 확정과 미확정을 섞으면 사건이 실제보다 더 커지거나 더 작아 보입니다.</p>

<h2 id="thakicloud-관점-데이터셋-처리를-신뢰-경계로-다루기">ThakiCloud 관점: 데이터셋 처리를 신뢰 경계로 다루기</h2>

<p>이번 사건이 인프라 회사에 주는 교훈은 분명합니다. 데이터셋은 수동적인 파일이 아니라, 처리되는 순간 코드를 실행할 수 있는 능동적 입력이라는 사실입니다. 그래서 우리는 이 문제를 두 개의 렌즈로 봅니다.</p>

<p><strong>ai-platform 렌즈에서</strong>, ThakiCloud의 ai-platform은 K8s 기반의 멀티테넌트 AI/ML 인프라입니다. 이런 환경에서 데이터셋 로딩과 전처리는 신뢰 경계 안쪽이 아니라 바깥쪽 입력으로 취급돼야 합니다. 구체적으로는 데이터셋 처리 작업을 권한이 최소화된 격리 컨테이너에서 실행하고, 네트워크 이그레스를 기본 차단하며, 노드 자격증명과 클라우드 크리덴셜을 워크로드가 직접 만질 수 없도록 분리하는 설계입니다. 이번 침해가 노드 레벨 접근에서 자격증명 탈취로 번졌다는 점은, 실행 격리와 크리덴셜 분리가 왜 옵션이 아니라 기본값이어야 하는지를 다시 보여줍니다. 온프렘과 주권 AI 수요가 높은 이유도 여기에 있습니다. 데이터와 실행이 고객 경계 안에 머무를수록 이런 파이프라인 공격의 폭발 반경을 줄일 수 있습니다.</p>

<p><strong>Paxis 렌즈에서</strong>, 이번 사건은 에이전트 네이티브 클라우드가 애초에 겨냥하는 위협 모델과 정확히 겹칩니다. Paxis는 ThakiCloud의 Agent-Native Cloud로, 스킬과 도구를 격리된 샌드박스에서 실행하고 모든 행동을 정책 게이트와 감사 로그로 통과시키는 것을 일급 원칙으로 삼습니다. 공격자가 자율 에이전트 스웜으로 수천 개의 액션을 던졌다는 점은, 에이전트의 행동을 실행 전에 정책으로 심사하고 실행 후에 감사 로그로 남기는 구조가 왜 필요한지를 그대로 증명합니다. 단기 샌드박스를 쓰고 버리는 공격 패턴에 맞서려면, 방어 측도 각 실행을 격리하고 그 실행의 권한 범위를 명시적으로 스코프하며 되돌릴 수 있는 감사 흔적을 남겨야 합니다. 격리 실행과 정책·감사는 에이전트 시대의 사치가 아니라 최소 요건입니다.</p>

<p>두 렌즈는 서로를 보완합니다. ai-platform이 데이터셋 처리라는 인프라 층에서 폭발 반경을 좁히고, Paxis가 에이전트 행동이라는 제어 층에서 각 액션을 심사합니다. 이번처럼 진입은 데이터 파이프라인이고 확산은 자율 에이전트인 공격에서는, 두 층의 방어가 함께 있어야 체인을 끊을 수 있습니다.</p>

<h2 id="한계-및-반론">한계 및 반론</h2>

<p>이 글의 결론을 과신하지 않도록 몇 가지를 분명히 해 둡니다. 첫째, 사건의 세부는 여전히 확정 중입니다. 정확한 액션 수, 자격증명 탈취의 범위, 상용 모델 협조 거부 일화 같은 색깔 있는 디테일은 이차 보도에 크게 의존하며, 공식 공지의 확정 사실과 구분해야 합니다.</p>

<p>둘째, 우리의 방어 서술이 곧 완결된 안전을 뜻하지는 않습니다. 격리와 정책·감사는 폭발 반경을 줄이는 설계 원칙이지, 취약점 자체를 없애는 마법이 아닙니다. 데이터셋 로더의 원격 코드 실행이나 설정 파싱의 인젝션 같은 취약점은 코드 수준에서 계속 발견되고 패치되어야 하며, 격리는 그 취약점이 터졌을 때 피해를 가두는 두 번째 방어선입니다.</p>

<p>셋째, 자율 에이전트 공격을 과대평가하는 것도 위험합니다. 이번 침해의 근본 원인은 정교한 AI가 아니라, 신뢰 경계를 넘는 입력이 코드를 실행할 수 있었던 익숙한 취약점 두 개였습니다. 에이전트는 그 취약점을 더 빠르고 넓게 악용하는 자동화였을 뿐입니다. 따라서 대응의 우선순위는 여전히 기본기에 있습니다. 신뢰할 수 없는 입력을 실행 권한과 분리하고, 자격증명을 워크로드에서 떼어 내며, 모든 실행을 관측 가능하게 만드는 일입니다.</p>

<p>허깅페이스의 신속한 봉쇄와 투명한 공개는 좋은 대응 사례로 남을 것입니다. 우리에게 남는 숙제는 단순합니다. 데이터셋을 파일이 아니라 코드로 대하는 것, 그리고 에이전트의 모든 행동을 심사와 감사의 대상으로 두는 것입니다.</p>

<h2 id="출처">출처</h2>

<ul>
  <li><a href="https://huggingface.co/blog/security-incident-july-2026">Security incident disclosure, July 2026 (Hugging Face 공식 블로그)</a></li>
  <li><a href="https://www.helpnetsecurity.com/2026/07/20/hugging-face-breached-by-autonomous-ai-agent/">Hugging Face breached by autonomous AI agent (Help Net Security)</a></li>
  <li><a href="https://www.bleepingcomputer.com/news/security/hugging-face-breach-autonomous-ai-agent-system-internal-datasets-credentials/">Hugging Face warns an autonomous AI agent hacked its network (BleepingComputer)</a></li>
  <li><a href="https://thehackernews.com/2026/07/worlds-largest-ai-model-repository.html">World’s Largest AI Model Repository Hugging Face Breached by Autonomous AI Agent (The Hacker News)</a></li>
  <li>이차 보도(정확한 액션 수·모델 협조 거부 일화 등은 확정 사실이 아닌 보도 인용): Cryptobriefing, Undercode Testing</li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="news" /><category term="security" /><category term="huggingface" /><category term="ai-agent" /><category term="supply-chain" /><category term="sandbox" /><category term="dataset-security" /><category term="news" /><category term="thakicloud" /><summary type="html"><![CDATA[허깅페이스가 2026년 7월 자율 AI 에이전트에 의한 내부 침해를 공개했습니다. 진입로는 악성 데이터셋 하나였고, 데이터셋 처리 파이프라인의 두 취약점이 코드 실행으로 이어졌습니다. 확인된 사실과 아직 조사 중인 부분을 나누고, 데이터셋 처리를 신뢰 경계로 다뤄야 하는 이유를 정리했습니다.]]></summary></entry><entry xml:lang="ko"><title type="html">동적 배치 튜닝이 정적 스케줄링을 이기지 못한 이유: 멀티테넌트 vLLM 서빙 제어루프의 정직한 실패 분석</title><link href="https://thakicloud.github.io/ko/research/agent-dynamic-batch-tuning-vllm/" rel="alternate" type="text/html" title="동적 배치 튜닝이 정적 스케줄링을 이기지 못한 이유: 멀티테넌트 vLLM 서빙 제어루프의 정직한 실패 분석" /><published>2026-07-21T00:00:00+09:00</published><updated>2026-07-21T00:00:00+09:00</updated><id>https://thakicloud.github.io/ko/research/agent-dynamic-batch-tuning-vllm</id><content type="html" xml:base="https://thakicloud.github.io/ko/research/agent-dynamic-batch-tuning-vllm/"><![CDATA[<p>Kubernetes 위에서 H200 같은 고사양 GPU를 여러 테넌트가 나눠 쓰는 vLLM 서빙 클러스터를 운영하는 분이라면, 그리고 Kueue의 정적 어드미션 제어를 “이 정도면 충분하다”고 여겨왔던 분이라면 이 글이 유용합니다. 반대로 “LLM 에이전트가 실시간으로 서빙 파라미터를 조정하면 무조건 더 낫다”고 가정하고 있었다면 더더욱 읽어볼 가치가 있습니다. 이 논문은 그 가정이 항상 참은 아니라는 것을, 그것도 꽤 구체적인 이유와 함께 보여주기 때문입니다.</p>

<h2 id="문제의식-어드미션은-정적인데-부하는-실시간이다">문제의식: 어드미션은 정적인데 부하는 실시간이다</h2>

<p>공유 GPU 위에서 여러 테넌트의 LLM 추론을 서빙하는 문제는 이제 흔한 운영 과제가 됩니다. H200 같은 대용량 가속기 한 장에 트래픽의 버스트성, 시퀀스 길이, SLO가 서로 다른 여러 테넌트가 동시에 올라탑니다. Kubernetes 기반 플랫폼에서는 보통 Kueue 같은 쿼터·큐 기반 스케줄러가 어떤 작업을 풀에 들여보낼지 결정하지만, 일단 들어온 요청들이 엔진의 유한한 용량을 실제로 어떻게 나눠 쓰는지를 좌우하는 배치 크기와 최대 동시 시퀀스 수 같은 서빙 시점 파라미터는 대개 배포 시점에 정적으로 고정됩니다.</p>

<p>이 논문은 여기서 자연스럽게 따라오는 질문을 던집니다. Kueue가 관리하는 멀티테넌트 H200 GPU 풀에서 여러 테넌트가 vLLM 추론 서버를 공유할 때, 실시간 GPU 메모리·지연 텔레메트리를 관측하는 LLM 에이전트가 테넌트별 배치 크기와 최대 동시 시퀀스 수를 온라인으로 재조정하면 정적 Kueue 어드미션 제어 대비 처리량, p99 지연, 비용의 파레토 프론티어를 실측으로 개선할 수 있을까요. 저자들은 이 질문에 답하기 위해 단일 H200에서 돌아가는 완전한 실측 프로토콜을 설계했지만, 목표 Kubernetes 클러스터 컨텍스트에 접근할 수 없어 실행 자체가 무산됐다는 사실을 서두에서부터 분명히 밝힙니다. 그래서 이 논문에는 실제 하드웨어에서 측정한 처리량이나 지연 수치가 단 하나도 등장하지 않습니다. 대신 논문은 세 가지를 내놓습니다. 온라인 배치·동시성 튜닝 문제의 제어루프 정식화, 클러스터 접근이 복구되는 즉시 실행 가능한 재현 가능 프로토콜, 그리고 세 가지 제어 법칙의 구조를 스트레스 테스트하는 결정론적 큐잉 시뮬레이션입니다.</p>

<h2 id="제어루프로-정식화한-튜닝-문제">제어루프로 정식화한 튜닝 문제</h2>

<p>논문은 온라인 테넌트별 배치·동시성 튜닝을 마르코프 결정 과정과 유사한 구조를 가진 이산 시간 폐루프 제어 문제로 정식화합니다. 상태는 테넌트별 큐 깊이, GPU 메모리 사용률, 최근 지연 백분위수(p50/p95/p99), 현재 동시성 상한으로 구성되고, 액션은 제한된 폭의 테넌트별 동시성 증감분입니다. 보상 함수는 처리량에서 SLO를 넘긴 p99 지연에 대한 페널티와 비용 항을 뺀 형태로, 안전 범위 안에서 처리량-지연-비용을 함께 최적화하도록 설계됩니다. 이 정식화 위에서 저자들은 에이전트가 실제 배포에 통합되는 방식도 함께 제안합니다. 에이전트는 Kueue의 잡 어드미션을 대체하지 않고 그 안쪽에서 사이드카처럼 동작하며, 이미 어드미션된 작업들이 엔진 용량을 나누는 방식만 조정합니다. 핵심은 손대는 레버가 vLLM 엔진 내부의 <code class="language-plaintext highlighter-rouge">max_num_seqs</code>가 아니라 클라이언트 측 테넌트별 어드미션 동시성이라는 점입니다. 프로덕션 vLLM은 엔진 파라미터를 재시작 없이 실시간으로 바꿀 수 있는 노브로 노출하지 않기 때문에, 실제로 조정 가능한 지점은 서버에 들어오는 동시 요청 수뿐이라는 실무적 제약을 그대로 반영한 설계입니다.</p>

<p><img src="/assets/images/posts/research/agent-dynamic-batch-tuning-vllm/fig-control-loop.png" alt="Control Loop: Agent-Driven Tuning Architecture" />
<em>LLM 에이전트가 GPU 텔레메트리와 테넌트별 큐 깊이를 관측해 제한된 폭의 동시성 조정값을 산출하고, 이를 통해 배치 구성을 재조정하는 제어루프 구조입니다. 개념적 아키텍처 다이어그램이며 하드웨어에서 실측된 결과가 아닙니다.</em></p>

<h2 id="시뮬레이션이-던진-경고-순진한-동적-제어는-정적보다-나빴다">시뮬레이션이 던진 경고: 순진한 동적 제어는 정적보다 나빴다</h2>

<p>실측 프로토콜이 실행되지 못한 자리를 메우기 위해 저자들은 Python 표준 라이브러리만으로 작성한 시드 고정 이산 시간 큐잉 시뮬레이션을 만듭니다. 두 테넌트가 유한한 슬롯 풀을 공유하는 단순화된 모델 위에서, 세 가지 정책을 20개 시드·1800초씩 돌려 평균을 냅니다. 정적 정책은 테넌트별 동시성 상한을 4로 고정하고, 순진한 동적 정책은 GPU 메모리 사용률과 6초 윈도우 p99 지연을 3초마다 관측해 두 테넌트를 한꺼번에 ±1씩 조절하며, 차별화 에이전트 정책은 테넌트별 백로그 비율을 각자 독립적으로 반영하는 방식으로 더 정교한 에이전트 추론을 흉내 낸 대리 모델입니다.</p>

<p>결과는 예상을 벗어납니다. 정적 정책이 처리량(0.686 req/s)과 드롭 수(평균 1.9건) 모두에서 가장 우수했고, 순진한 동적 정책은 처리량이 0.515 req/s로 떨어지면서 평균 309.5건이나 요청을 드롭했습니다. 차별화 에이전트 대리 모델은 순진한 동적 정책보다는 나았지만(처리량 0.601 req/s, 드롭 154.7건) 여전히 정적 정책을 따라잡지 못했고, p99 지연 역시 두 동적 변형 모두 정적 정책(11.75초)보다 높게 나타났습니다(각각 12.79초, 13.25초).</p>

<p><img src="/assets/images/posts/research/agent-dynamic-batch-tuning-vllm/fig-throughput-dropped.png" alt="Throughput vs. Dropped Requests by Policy" />
<em>정적 정책이 가장 높은 처리량과 가장 적은 드롭을 기록했고, 두 동적 변형 모두 시뮬레이션에서 성능이 낮았습니다. 20개 시드 평균을 낸 결정론적 큐잉 시뮬레이션 결과이며 실제 GPU에서 측정한 값이 아닙니다.</em></p>

<p><img src="/assets/images/posts/research/agent-dynamic-batch-tuning-vllm/fig-p99-latency.png" alt="p99 Latency by Policy" />
<em>정적 정책의 p99 지연이 가장 낮았고, 두 동적 컨트롤러 모두 이 시뮬레이션 영역에서 꼬리 지연을 줄이지 못했습니다. 20개 시드 평균의 시뮬레이션 결과이며 하드웨어 실측치가 아닙니다.</em></p>

<p>저자들은 이 결과를 단순 버그가 아니라 모델의 실제 동역학이라고 검증하며 네 단계로 원인을 추적합니다. 첫째, 서비스 시간이 지수분포를 따르기 때문에 원래부터 꼬리가 두껍습니다. 평균 2.5초짜리 지수분포는 큐잉이 전혀 없어도 p99가 11.5초 근처에 형성되는데, 이는 정적 정책이 기록한 11.75초와 거의 차이가 없습니다. 즉 관측되는 꼬리 지연 대부분은 혼잡이 아니라 서비스 시간 자체의 분산에서 나옵니다. 둘째, 스로틀 임계값이 3초 기준의 두 배인 6초로 설정되어 있는데, 이는 이 내재적 p99인 11.5초보다 훨씬 낮습니다. 몇 건 안 되는 완료 요청으로 추정하는 짧은 윈도우 p99는 노이즈가 크고 이 내재적 꼬리 쪽으로 편향되어, 실제로는 가볍게 부하가 걸린 상황에서도 임계값을 자주 넘깁니다. 셋째, 이렇게 발생한 오탐 스로틀이 두 테넌트를 함께 묶어 낮추기 때문에 혼잡하지 않은 테넌트까지 함께 제한받고, 3초짜리 하드 타임아웃 아래서 불필요한 스로틀 하나하나가 곧바로 드롭으로 전환됩니다. 넷째, 확장 규칙으로 회복은 되지만 스로틀-드롭 비용이 확장 이득보다 비대칭적으로 크기 때문에 순 처리량이 떨어집니다.</p>

<h2 id="실패에서-뽑아낸-설계-요구사항">실패에서 뽑아낸 설계 요구사항</h2>

<p>이 진단이 말해주는 바는 동적 튜닝 자체가 쓸모없다는 것이 아닙니다. 노이즈가 크고 꼬리 쪽으로 편향된 신호, 테넌트를 함께 묶는 스로틀링, 하드 타임아웃이 결합하면 순진한 고정 임계값 제어가 정적 제어보다 실제로 더 나빠질 수 있다는 사실입니다. 여기서 저자들은 효과적인 에이전트 컨트롤러가 갖춰야 할 네 가지 설계 요구사항을 도출합니다. 혼잡 신호는 짧은 윈도우의 원시 백분위수 대신 더 긴 텔레메트리 이력과 신뢰도 추정을 바탕으로 견고하게 추정해야 합니다. 결정은 테넌트를 하나로 묶은 전역 신호가 아니라 테넌트별로 차별화되어야 합니다. 스로틀링과 확장 사이의 비대칭적 비용을 명시적으로 고려해야 하고, 고정된 숫자 임계값으로는 표현할 수 없는 버스트 패턴 인식이나 테넌트별 SLA 메타데이터 같은 추가 맥락도 활용해야 합니다. 실제로 차별화 에이전트 대리 모델이 순진한 컨트롤러 대비 손실의 절반가량을 회복한 것은 이 방향이 유효할 가능성을 시사합니다. 다만 결합도와 신호 종류를 동시에 바꿨기 때문에 어느 쪽이 회복에 기여했는지는 이 실험만으로 분리해낼 수 없다고 저자들은 선을 긋습니다.</p>

<h2 id="회사사회과학에-남기는-것">회사·사회·과학에 남기는 것</h2>

<p>ThakiCloud 입장에서 이 연구가 당장 남기는 것은 검증된 비용 절감이 아닙니다. 우리 멀티테넌트 H200 클러스터에 곧바로 적용 가능한 재현 가능 실측 프로토콜, 그리고 순진하게 접근하면 오히려 손해를 볼 수 있다는 구체적 경고입니다. 클러스터 접근이 복구되면 이 프로토콜에 실제 LLM 에이전트 결정 함수만 대체해 넣으면 되도록 하네스 자체는 이미 완성돼 있습니다. 더 넓게 보면 공유 GPU 풀의 자원 낭비를 줄이는 방법론이 다듬어질수록 소규모 조직도 공유 클러스터에서 SLA를 지킬 수 있는 실용적 경로가 열립니다. 이는 추론 인프라의 에너지·비용 효율 전반에 도움이 됩니다. 과학적으로는 LLM 에이전트가 실시간 시스템 텔레메트리를 관측해 서빙 하이퍼파라미터를 온라인으로 재조정하는 제어루프를 명시적으로 정식화하고, 그 실패 모드를 재현 가능한 시뮬레이션으로 정량 진단했다는 점이 기여입니다. 이 논문은 ThakiCloud가 같은 Kueue·GPU 기반 위에서 이어온 연구 계보의 일부이기도 합니다. 저지 모델 샘플링 예산을 스케줄링하는 ABJ-Gate, 어드미션 시점의 TEE 원격 증명을 다루는 Attested Confidential Sovereign Inference, 드문 이산 사고에 대한 이진 원격·조치 결정을 다루는 “Escalate or Act?”와 비교하면 이 논문은 초 단위로 연속 재조정되는 실시간 다목적 제어라는 질적으로 다른 문제를 다룬다는 점에서 이들과 구별됩니다.</p>

<h2 id="한계">한계</h2>

<p>이 논문이 스스로 밝히는 한계는 명확합니다. 가장 중요한 것은 어떤 형태로든 실제 하드웨어 검증이 없다는 점으로, 목표 클러스터 컨텍스트에 접근할 수 없어 실측 프로토콜이 실행되지 못했고 논문에 실린 어떤 수치도 GPU에서 측정된 것이 아닙니다. 시뮬레이션은 토큰과 KV 캐시를 단일 스칼라 슬롯 수로 단순화했을 뿐 프리필/디코드나 메모리 압력 동역학을 모델링하지 않으며, 메모리 사용률 신호조차 슬롯 점유율의 대리값일 뿐 실제 디바이스 메모리 측정치가 아닙니다. 테스트한 테넌트 유형도 두 가지 합성 트래픽 패턴에 그쳐 다른 트래픽 혼합에서는 발견된 실패 모드가 일반화되지 않을 수 있습니다. 시뮬레이션 속 에이전트 정책 역시 실제 LLM이 아닌 손으로 작성한 대리 모델이며 결합도와 신호 종류를 동시에 바꿨기 때문에, 4장에서 도출한 설계 요구사항을 실제 LLM 에이전트가 충족한다는 검증된 근거는 아직 없습니다. 20개 시드에 대한 평균과 표준편차만 보고했을 뿐 통계적 유의성 검정은 수행하지 않았다는 점도 저자들은 분명히 밝힙니다. 이런 이유로 이 논문은 검증된 결과가 아니라 정식화와 프로토콜, 그리고 경고성 시뮬레이션으로서 스스로를 자리매김합니다.</p>

<p>논문 상세 페이지는 다음에서 확인할 수 있습니다: <a href="https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-21-agent-dynamic-batch-tuning-vllm">https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-21-agent-dynamic-batch-tuning-vllm</a></p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;nil, &quot;bio&quot;=&gt;nil, &quot;location&quot;=&gt;&quot;Seoul, Korea&quot;, &quot;email&quot;=&gt;&quot;info@thakicloud.co.kr&quot;, &quot;uri&quot;=&gt;nil, &quot;home&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;, &quot;url&quot;=&gt;&quot;https://thakicloud.co.kr&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/company/thakicloud&quot;}, {&quot;label&quot;=&gt;&quot;X&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-x-twitter&quot;, &quot;url&quot;=&gt;&quot;https://x.com/thakicloud&quot;}]}</name><email>info@thakicloud.co.kr</email></author><category term="research" /><category term="vLLM" /><category term="Kueue" /><category term="GPU 스케줄링" /><category term="멀티테넌트 서빙" /><category term="LLM 에이전트" /><category term="동적 배치" /><category term="추론 비용 최적화" /><category term="H200" /><category term="큐잉 시뮬레이션" /><category term="제어루프" /><summary type="html"><![CDATA[GPU 텔레메트리를 관측해 배치와 동시성을 실시간으로 조정하는 에이전트는 정적 Kueue 어드미션보다 나을까요. 시뮬레이션은 아니라고 답했고, 왜 그런지까지 원인을 추적했습니다.]]></summary></entry></feed>