<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Get started on Modelplane Docs</title><link>/getting-started/</link><description>Recent content in Get started on Modelplane Docs</description><generator>Hugo -- gohugo.io</generator><language>en-us</language><lastBuildDate>Mon, 01 Jan 0001 00:00:00 +0000</lastBuildDate><atom:link href="/getting-started/index.xml" rel="self" type="application/rss+xml"/><item><title>Build the platform</title><link>/getting-started/build-the-platform/</link><pubDate/><guid>/getting-started/build-the-platform/</guid><description>&lt;p&gt;This is the platform team&amp;rsquo;s side of Modelplane. You set up the gateway that
fronts your models, give the control plane cloud credentials, and register your
first GPU cluster: a hardware profile published as an &lt;code&gt;InferenceClass&lt;/code&gt; and an
&lt;code&gt;InferenceCluster&lt;/code&gt; that offers it.&lt;/p&gt;
&lt;p&gt;In the next step, the ML team will create a model deployment that schedules
against this capacity without knowing which cluster it runs on.&lt;/p&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites &lt;a class="anchor-link" id="prerequisites" href="#prerequisites" aria-label="Link to this section: Prerequisites"&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;ul class="nav nav-tabs" id="644" role="tablist"&gt;
&lt;li class="nav-item" role="presentation"&gt;
&lt;button class="nav-link active "
id="tab-1439"
data-bs-toggle="tab"
data-bs-target="#tab-pane-1439"
type="button"
role="tab"
aria-controls="tab-pane-1439" aria-selected="false"&gt;EKS&lt;/button&gt;
&lt;/li&gt;
&lt;li class="nav-item" role="presentation"&gt;
&lt;button class="nav-link "
id="tab-1505"
data-bs-toggle="tab"
data-bs-target="#tab-pane-1505"
type="button"
role="tab"
aria-controls="tab-pane-1505" aria-selected="false"&gt;GKE&lt;/button&gt;
&lt;/li&gt;
&lt;li class="nav-item" role="presentation"&gt;
&lt;button class="nav-link "
id="tab-283"
data-bs-toggle="tab"
data-bs-target="#tab-pane-283"
type="button"
role="tab"
aria-controls="tab-pane-283" aria-selected="false"&gt;AKS&lt;/button&gt;
&lt;/li&gt;
&lt;li class="nav-item" role="presentation"&gt;
&lt;button class="nav-link "
id="tab-1263"
data-bs-toggle="tab"
data-bs-target="#tab-pane-1263"
type="button"
role="tab"
aria-controls="tab-pane-1263" aria-selected="false"&gt;Nebius&lt;/button&gt;
&lt;/li&gt;
&lt;li class="nav-item" role="presentation"&gt;
&lt;button class="nav-link "
id="tab-1546"
data-bs-toggle="tab"
data-bs-target="#tab-pane-1546"
type="button"
role="tab"
aria-controls="tab-pane-1546" aria-selected="false"&gt;Vultr&lt;/button&gt;
&lt;/li&gt;
&lt;li class="nav-item" role="presentation"&gt;
&lt;button class="nav-link "
id="tab-520"
data-bs-toggle="tab"
data-bs-target="#tab-pane-520"
type="button"
role="tab"
aria-controls="tab-pane-520" aria-selected="false"&gt;Civo&lt;/button&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="tab-content" id="Content-644"&gt;
&lt;div
class="tab-pane fade active show pt-sm-3"
id="tab-pane-1439"
role="tabpanel"
aria-labelledby="tab-1439"
tabindex="0"&gt;&lt;ul&gt;
&lt;li&gt;An AWS account with permissions to create EKS clusters, VPCs, and IAM roles&lt;/li&gt;
&lt;li&gt;AWS access key ID and secret access key&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;
&lt;div
class="tab-pane fade pt-sm-3"
id="tab-pane-1505"
role="tabpanel"
aria-labelledby="tab-1505"
tabindex="0"&gt;&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;A GCP service account JSON key, granted these roles on the project:&lt;/p&gt;</description></item><item><title>Deploying a model</title><link>/getting-started/deploying-a-model/</link><pubDate/><guid>/getting-started/deploying-a-model/</guid><description>&lt;p&gt;Now that the platform is provisioned, the ML team can declare what a model needs
with a &lt;code&gt;ModelDeployment&lt;/code&gt;. Describe the hardware requirements and the scheduler
schedules against the capacity the platform team published.&lt;/p&gt;
&lt;h2 id="create-a-deployment"&gt;Create a deployment &lt;a class="anchor-link" id="create-a-deployment" href="#create-a-deployment" aria-label="Link to this section: Create a deployment"&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Create a namespace for the model:&lt;/p&gt;
&lt;div class="code-card"&gt;
&lt;div class="code-card__header"&gt;
&lt;span class="code-card__name"&gt;bash&lt;/span&gt;
&lt;div class="code-card__actions"&gt;
&lt;button class="code-card__btn code-card__copy" type="button" aria-label="Copy contents" title="Copy contents"&gt;
&lt;svg class="icon-main" width="15" height="15" viewBox="0 0 16 16" fill="currentColor" aria-hidden="true" focusable="false"&gt;&lt;path d="M4 1.5H3a2 2 0 0 0-2 2V14a2 2 0 0 0 2 2h10a2 2 0 0 0 2-2V3.5a2 2 0 0 0-2-2h-1v1h1a1 1 0 0 1 1 1V14a1 1 0 0 1-1 1H3a1 1 0 0 1-1-1V3.5a1 1 0 0 1 1-1h1z"/&gt;&lt;path d="M9.5 1a.5.5 0 0 1 .5.5v1a.5.5 0 0 1-.5.5h-3a.5.5 0 0 1-.5-.5v-1a.5.5 0 0 1 .5-.5zm-3-1A1.5 1.5 0 0 0 5 1.5v1A1.5 1.5 0 0 0 6.5 4h3A1.5 1.5 0 0 0 11 2.5v-1A1.5 1.5 0 0 0 9.5 0z"/&gt;&lt;/svg&gt;
&lt;svg class="icon-check" width="16" height="16" viewBox="0 0 16 16" fill="currentColor" aria-hidden="true" focusable="false"&gt;&lt;path d="M13.485 1.929a.75.75 0 0 1 .086 1.057l-7 8.5a.75.75 0 0 1-1.117.063l-3.5-3.5a.75.75 0 0 1 1.06-1.06l2.92 2.92 6.494-7.894a.75.75 0 0 1 1.057-.086z"/&gt;&lt;/svg&gt;
&lt;/button&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div class="code-card__body"&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;kubectl create namespace ml-team&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;The device selector matches against the capacity declared in the
&lt;code&gt;InferenceClass&lt;/code&gt;, not the pod&amp;rsquo;s resource requests. Any L4 node satisfies
&lt;code&gt;&amp;gt;= 20Gi&lt;/code&gt;, so this deployment runs on the cluster you just added:&lt;/p&gt;</description></item><item><title>Scale the platform</title><link>/getting-started/scale-the-platform/</link><pubDate/><guid>/getting-started/scale-the-platform/</guid><description>&lt;p&gt;You have one small-GPU cluster with a running model. In this guide, you&amp;rsquo;ll grow
the fleet with larger-GPU capacity so the ML team has more to schedule against.&lt;/p&gt;
&lt;p&gt;Provisioning takes about 10 to 15 minutes.&lt;/p&gt;
&lt;h2 id="register-more-clusters"&gt;Register more clusters &lt;a class="anchor-link" id="register-more-clusters" href="#register-more-clusters" aria-label="Link to this section: Register more clusters"&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;ul class="nav nav-tabs" id="556" role="tablist"&gt;
&lt;li class="nav-item" role="presentation"&gt;
&lt;button class="nav-link active "
id="tab-400"
data-bs-toggle="tab"
data-bs-target="#tab-pane-400"
type="button"
role="tab"
aria-controls="tab-pane-400" aria-selected="false"&gt;EKS&lt;/button&gt;
&lt;/li&gt;
&lt;li class="nav-item" role="presentation"&gt;
&lt;button class="nav-link "
id="tab-1953"
data-bs-toggle="tab"
data-bs-target="#tab-pane-1953"
type="button"
role="tab"
aria-controls="tab-pane-1953" aria-selected="false"&gt;GKE&lt;/button&gt;
&lt;/li&gt;
&lt;li class="nav-item" role="presentation"&gt;
&lt;button class="nav-link "
id="tab-1468"
data-bs-toggle="tab"
data-bs-target="#tab-pane-1468"
type="button"
role="tab"
aria-controls="tab-pane-1468" aria-selected="false"&gt;AKS&lt;/button&gt;
&lt;/li&gt;
&lt;li class="nav-item" role="presentation"&gt;
&lt;button class="nav-link "
id="tab-1156"
data-bs-toggle="tab"
data-bs-target="#tab-pane-1156"
type="button"
role="tab"
aria-controls="tab-pane-1156" aria-selected="false"&gt;Nebius&lt;/button&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="tab-content" id="Content-556"&gt;
&lt;div
class="tab-pane fade active show pt-sm-3"
id="tab-pane-400"
role="tabpanel"
aria-labelledby="tab-400"
tabindex="0"&gt;&lt;p&gt;Register two more clusters with a bigger hardware class: &lt;code&gt;L40S&lt;/code&gt; (&lt;code&gt;48 GB&lt;/code&gt;) in
&lt;code&gt;us-west&lt;/code&gt; and &lt;code&gt;eu-central&lt;/code&gt;:&lt;/p&gt;</description></item><item><title>Scale the model</title><link>/getting-started/scale-the-model/</link><pubDate/><guid>/getting-started/scale-the-model/</guid><description>&lt;p&gt;A &lt;code&gt;ModelService&lt;/code&gt; can front more than one &lt;code&gt;ModelDeployment&lt;/code&gt;. Here you add a second
deployment, pinned to a different region, and point the same service at both. The
endpoint you already curled stays the same. Behind it, traffic now load-balances
across two regions.&lt;/p&gt;
&lt;pre class="mermaid"&gt;graph LR
subgraph fleet ["Fleet"]
IC1["cluster-a\nsmall GPU"]
IC2["cluster-b\nlarger GPU"]
end
subgraph ml ["ML team"]
MD1["ModelDeployment\nqwen-demo"]
MD2["ModelDeployment\nqwen-west\ntargets cluster-b"]
MS["ModelService qwen\nmodel: ml-team/qwen"]
end
IC1 --&gt; MD1
IC2 --&gt; MD2
MD1 --&gt; MS
MD2 --&gt; MS
&lt;/pre&gt;
&lt;h2 id="add-a-second-deployment"&gt;Add a second deployment &lt;a class="anchor-link" id="add-a-second-deployment" href="#add-a-second-deployment" aria-label="Link to this section: Add a second deployment"&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The new deployment targets a larger-GPU cluster you added in the last step. On
EKS, GKE, and AKS it pins to a second region with a &lt;code&gt;clusterSelector&lt;/code&gt;. On Nebius,
which runs one region per project, the capacity selector alone routes it to the
&lt;code&gt;H100&lt;/code&gt; tier:&lt;/p&gt;</description></item><item><title>Clean up</title><link>/getting-started/clean-up/</link><pubDate/><guid>/getting-started/clean-up/</guid><description>&lt;p&gt;Delete the model resources, the gateway, the clusters, and finally the control
plane.&lt;/p&gt;
&lt;h2 id="delete-model-resources"&gt;Delete model resources &lt;a class="anchor-link" id="delete-model-resources" href="#delete-model-resources" aria-label="Link to this section: Delete model resources"&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Delete model resources before clusters. A cluster refuses deletion while
anything still runs on it. Foreground cascading deletion holds each resource
until what it composed on the clusters is gone, so a cluster isn&amp;rsquo;t released
while that&amp;rsquo;s still being removed:&lt;/p&gt;
&lt;div class="code-card"&gt;
&lt;div class="code-card__header"&gt;
&lt;span class="code-card__name"&gt;bash&lt;/span&gt;
&lt;div class="code-card__actions"&gt;
&lt;button class="code-card__btn code-card__copy" type="button" aria-label="Copy contents" title="Copy contents"&gt;
&lt;svg class="icon-main" width="15" height="15" viewBox="0 0 16 16" fill="currentColor" aria-hidden="true" focusable="false"&gt;&lt;path d="M4 1.5H3a2 2 0 0 0-2 2V14a2 2 0 0 0 2 2h10a2 2 0 0 0 2-2V3.5a2 2 0 0 0-2-2h-1v1h1a1 1 0 0 1 1 1V14a1 1 0 0 1-1 1H3a1 1 0 0 1-1-1V3.5a1 1 0 0 1 1-1h1z"/&gt;&lt;path d="M9.5 1a.5.5 0 0 1 .5.5v1a.5.5 0 0 1-.5.5h-3a.5.5 0 0 1-.5-.5v-1a.5.5 0 0 1 .5-.5zm-3-1A1.5 1.5 0 0 0 5 1.5v1A1.5 1.5 0 0 0 6.5 4h3A1.5 1.5 0 0 0 11 2.5v-1A1.5 1.5 0 0 0 9.5 0z"/&gt;&lt;/svg&gt;
&lt;svg class="icon-check" width="16" height="16" viewBox="0 0 16 16" fill="currentColor" aria-hidden="true" focusable="false"&gt;&lt;path d="M13.485 1.929a.75.75 0 0 1 .086 1.057l-7 8.5a.75.75 0 0 1-1.117.063l-3.5-3.5a.75.75 0 0 1 1.06-1.06l2.92 2.92 6.494-7.894a.75.75 0 0 1 1.057-.086z"/&gt;&lt;/svg&gt;
&lt;/button&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div class="code-card__body"&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;kubectl delete md --all -n ml-team --cascade&lt;span class="o"&gt;=&lt;/span&gt;foreground
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;kubectl delete ms --all -n ml-team --cascade&lt;span class="o"&gt;=&lt;/span&gt;foreground&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="delete-the-gateway"&gt;Delete the gateway &lt;a class="anchor-link" id="delete-the-gateway" href="#delete-the-gateway" aria-label="Link to this section: Delete the gateway"&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Delete the gateway before its cluster. The &lt;code&gt;InferenceGateway&lt;/code&gt; runs a load balancer
on the cluster it names; deleting it while that cluster is still up lets the load
balancer be removed, rather than leaking it when the cluster goes. Foreground
deletion holds the gateway until its objects on the cluster are deleted:&lt;/p&gt;</description></item></channel></rss>