<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>VLLM on G. T. Wang</title>
    <link>https://blog.gtwang.org/tags/vllm/</link>
    <description>Recent content in VLLM on G. T. Wang</description>
    <generator>Hugo -- 0.162.1</generator>
    <language>zh-tw</language>
    <copyright>G. T. Wang</copyright>
    <lastBuildDate>Wed, 09 Sep 2026 21:50:13 +0800</lastBuildDate>
    <atom:link href="https://blog.gtwang.org/tags/vllm/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>vLLM 高效能 LLM 推論框架安裝與使用教學</title>
      <link>https://blog.gtwang.org/linux/vllm-high-throughput-memory-efficient-llm-engine-setup-tutorial-20260909/</link>
      <pubDate>Wed, 09 Sep 2026 21:50:13 +0800</pubDate>
      <guid>https://blog.gtwang.org/linux/vllm-high-throughput-memory-efficient-llm-engine-setup-tutorial-20260909/</guid>
      <description>&lt;p&gt;本文介紹如何在 Ubuntu Linux 中安裝、設定與使用 vLLM 大型語言模型推論框架，安裝各種本地端模型，進行正式環境的效能調校與布署。&lt;/p&gt;
&lt;p&gt;

&lt;ins class=&#34;adsbygoogle&#34;
     style=&#34;display:block&#34;
     data-ad-client=&#34;ca-pub-7794009487786811&#34;
     data-ad-slot=&#34;9921134032&#34;
     data-ad-format=&#34;auto&#34;
     data-full-width-responsive=&#34;true&#34;&gt;&lt;/ins&gt;
&lt;script&gt;
     (adsbygoogle = window.adsbygoogle || []).push({});
&lt;/script&gt;
&lt;/p&gt;

&lt;p&gt;&lt;a href=&#34;https://vllm.ai/&#34;&gt;vLLM&lt;/a&gt; 是一套開放原始碼的 LLM 推理服務引擎框架，使用方式簡單、速度快、吞吐量高、延遲低，適合用於大規模運算的的正式服務環境，可有效利用 GPU 的所有運算能力。&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
