深色模式
模型服务 KServe 实战
摘要:本文面向需要在 Kubernetes 上以声明式、可灰度、可回滚方式部署模型的 SRE / 平台工程师。覆盖 KServe
InferenceService的核心结构、RawDeployment 与 Serverless 两种模式、sklearn 与 vLLM(LLM) 两类部署、金丝雀流量切分(canaryTrafficPercent)、验证与回滚。适用版本:KServe v0.20([版本相关:以官方发布为准]),Kubernetes v1.28+,apiVersionserving.kserve.io/v1beta1。
适用版本与前提
- Kubernetes:v1.28+(KServe Serverless 模式依赖 Knative Serving;RawDeployment 模式仅依赖标准 Deployment/Service)
- 已安装 KServe(Helm
kserve/kserve),本文示例以 RawDeployment 模式为主,避免引入 Knative 依赖 - 对象存储可用(S3 / MinIO / GCS)或 Hugging Face Hub 作为模型存储后端
- 如需 GPU 推理,节点已部署 NVIDIA GPU Operator(见 gpu-scheduling.md)
核心概念:InferenceService 的组成
一个 InferenceService 由三部分组成:
| 组件 | 必选 | 作用 |
|---|---|---|
predictor | 是 | 加载模型并对外提供推理(可用内置运行时或自定义容器) |
transformer | 否 | 请求预处理 / 响应后处理 |
explainer | 否 | 可解释性(如集成 Alibi) |
KServe 支持两种运行模式:
- Serverless(默认):底层用 Knative Serving,支持缩容到零、基于请求的 KPA 扩缩。
- RawDeployment:用原生 Deployment + Service + Istio/Ingress 暴露,运维更直观,本文示例使用此模式。
生产实践 1:sklearn 推理服务(RawDeployment)
yaml
# kserve-sklearn.yaml
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
name: sklearn-iris
namespace: models
annotations:
serving.kserve.io/deploymentMode: RawDeployment
spec:
predictor:
model:
modelFormat:
name: sklearn
storageUri: "gs://kfserving-examples/models/sklearn/1.0/model"
resources:
requests:
cpu: "1"
memory: 2Gi
limits:
cpu: "2"
memory: 4Gi1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
存储 URI 与凭证
storageUri 支持 gs://、s3://、hf://(Hugging Face)、oci:// 等。S3 凭证通过 Secret 注入,命名需符合 KServe storage initializer 约定(如 aws-secret)。切勿把 AK/SK 写进 YAML 或镜像。
应用并查看状态:
bash
kubectl apply -f kserve-sklearn.yaml
kubectl get inferenceservice sklearn-iris -n models
# 期望输出(示意,字段随版本略有差异):
# NAME URL READY
# sklearn-iris http://sklearn-iris.models.example.com True1
2
3
4
5
2
3
4
5
生产实践 2:vLLM 部署 LLM(GPU)
RawDeployment 下可用自定义容器挂载 vLLM 运行时,对外暴露 OpenAI 兼容接口:
yaml
# kserve-vllm.yaml
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
name: qwen2-5-7b
namespace: models
annotations:
serving.kserve.io/deploymentMode: RawDeployment
serving.kserve.io/enable-prometheus-scraping: "true"
spec:
predictor:
minReplicas: 1
maxReplicas: 4
containers:
- name: vllm
image: vllm/vllm-openai:v0.6.3 # [版本相关: 以 vLLM 官方 tag 为准]
args:
- "--model"
- "Qwen/Qwen2.5-7B-Instruct"
- "--tensor-parallel-size"
- "1"
- "--gpu-memory-utilization"
- "0.90"
ports:
- containerPort: 8000
resources:
limits:
nvidia.com/gpu: 1
memory: "24Gi"
cpu: "8"
requests:
nvidia.com/gpu: 1
memory: "20Gi"
cpu: "4"
readinessProbe:
httpGet:
path: /health
port: 8000
initialDelaySeconds: 60
periodSeconds: 10
env:
- name: HUGGING_FACE_HUB_TOKEN
valueFrom:
secretKeyRef:
name: hf-token
key: token1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
vLLM 版本与镜像
vLLM 镜像 tag 随发布频繁变化,本文示例 v0.6.3 仅为示意,请以 vLLM 官方-release tag 为准;大模型权重体积大,首次冷启动拉取权重可能数分钟,探针 initialDelaySeconds 要留足。模型名需与 HF Hub 上完全一致。
生产实践 3:金丝雀流量切分
KServe 通过 canaryTrafficPercent 在新旧版本间按权重分流,旧版本存储 URI 保留即实现"秒级回滚":
yaml
# kserve-canary.yaml —— 在 predictor 同级设置
spec:
predictor:
canaryTrafficPercent: 10 # 10% 流量打向新版本
model:
modelFormat:
name: sklearn
storageUri: "s3://ml-models/iris/v2/" # 新版本产物1
2
3
4
5
6
7
8
2
3
4
5
6
7
8
逐步放量(patch 即可,无需重建):
bash
kubectl patch inferenceservice sklearn-iris -n models \
--type='json' \
-p='[{"op":"replace","path":"/spec/predictor/canaryTrafficPercent","value":50}]'1
2
3
2
3
验证无误后移除 canary 字段,流量全部切到新版本,旧副本自动缩容:
bash
kubectl patch inferenceservice sklearn-iris -n models \
--type='json' \
-p='[{"op":"remove","path":"/spec/predictor/canaryTrafficPercent"}]'1
2
3
2
3
验证
bash
# sklearn 推理请求(V1 协议)
curl -X POST http://sklearn-iris.models.example.com/v1/models/sklearn-iris:predict \
-H "Content-Type: application/json" \
-d '{"instances": [[6.8, 2.8, 4.8, 1.4]]}'
# vLLM OpenAI 兼容请求
curl http://qwen2-5-7b.models.example.com/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"Qwen/Qwen2.5-7B-Instruct","messages":[{"role":"user","content":"你好"}]}'1
2
3
4
5
6
7
8
9
2
3
4
5
6
7
8
9
期望结果
返回 JSON 预测或 completion 即成功。若返回 503/超时,先查 Pod 状态与 storageUri 可访问性(最常见故障见下)。命令输出为结构示意,实际响应以模型为准,[未实测具体 token 输出]。
回滚与清理
bash
# 回滚:将 canaryTrafficPercent 设为 0,流量全部回到上一稳定版本
kubectl patch inferenceservice sklearn-iris -n models \
--type='json' \
-p='[{"op":"replace","path":"/spec/predictor/canaryTrafficPercent","value":0}]'
# 删除整个 InferenceService(注意:会删除底层 Deployment/Service)
kubectl delete inferenceservice sklearn-iris -n models1
2
3
4
5
6
7
2
3
4
5
6
7
删除即停服
删除 InferenceService 会级联删除对应负载,生产环境先确认流量已切走或已灰度完成;建议通过 GitOps(Argo CD)以 git revert 方式回滚,保留审计。
故障排查
| 现象 | 可能原因 | 排查 |
|---|---|---|
Pod 一直 Initializing | storageUri 不可达 / 凭证缺失 | kubectl logs <pod> -c storage-initializer 看拉取日志 |
| 推理 503 | 就绪探针未过 / 副本为 0 | kubectl get pods -n models、kubectl describe 探针 |
| GPU 0 设备 | 节点无 GPU 或 taint 未容忍 | kubectl describe node 查 nvidia.com/gpu 与 taint |
| 金丝雀不生效 | 非 Serverless 模式或字段位置错误 | 确认 canaryTrafficPercent 在 spec.predictor 下,且模式支持 |
安全与合规
推理服务的越权与数据泄露
- 网关鉴权:InferenceService 公开即意味着任何人可调用(含按 token 计费的 LLM)。务必在 Ingress/Istio 层加认证(oauth2-proxy / mTLS)与 WAF。
- 模型仓库凭证:
storageUri拉取权重的 Secret 仅授权给推理命名空间,禁止跨命名空间复用同一长期 AK。 - HF token 最小权限:挂载的
HUGGING_FACE_HUB_TOKEN应仅具备目标仓库 read 权限,避免全账号 token 泄露导致权重被改。 - egress 收敛:推理 Pod 不应随意出公网,用 NetworkPolicy 限制,敏感场景启用离线模式防止数据外传。
成本与性能
- GPU 占用:7B 级模型(如 Qwen2.5-7B,FP16 约 14GB 权重)+ KV Cache,单卡 24Gi 显存(如 L4/A10G)通常可承载;70B 级需多卡张量并行(tensor-parallel-size>1)或量化。
- 成本示例([价格随云厂商/时点变化,未实测]):单张 L4(24Gi)约 0.5–1.5 USD/GPU·小时;若
minReplicas长期为 1,单月约 360–1080 USD。通过maxReplicas限制峰值、低峰缩容到minReplicas控制费用。 - 吞吐:LLM 关注 token/s 与 TTFT(首 token 延迟),用 vLLM Prometheus 指标
vllm:...与 KServe 队列指标监控,[未实测具体数值]。