技能低风险未认领
AKS GPU Inference Day-2
Diagnose Day-2 AKS GPU and KAITO incidents using profile-aware, read-only evidence. WHEN: 'Insufficient nvidia.com/gpu', GPU pod Pending, model-load OOM, DCGM/VRAM, KAITO Workspace not ready, or GPU autoscaling. DO NOT USE FOR: setup (airunway-aks-setup), non-GPU incidents (aks-troubleshooting), standalone VM quota (azure-quotas), or generic cost (cost-analysis or cost-optimization from the optional azure-cost plugin).
microsoftmicrosoft/aks-gpu-inference
说明
Quick Reference
| Property | Value |
|---|---|
| Best for | AKS GPU/KAITO failures |
| Evidence | Bound target, events/status, gpuProfile, DCGM |
When to Use This Skill
Use for scheduling, VRAM/OOM, KAITO readiness, and workload scaling. Route exclusions as described above.
MCP Tools
Use a fitting host-advertised Azure read or the references' read-only queries.
Host Capability Gate
Before executing any referenced pipeline, verify that the host authorizes
the required shell and kubectl/az commands against the bound target.
File-backed collection also requires approved artifact storage and access to
any bundled files it uses. A governed Azure CLI tool alone does not establish
these capabilities. Equivalent host reads may replace commands only where
their advertised schemas provide the required evidence.
If execution is unavailable or prohibited, say so and analyze supplied or
redacted events, status, logs, and metrics, or give the operator a scoped
collection plan. State which reads did not run and leave conclusions requiring
missing evidence unconfirmed. Never route kubectl through Azure MCP or
bypass host policy. The authorization requirements below still apply.
Workflow
- Bind subscription, cluster, kube context, pool, and affected resource.
- Capture exact events/status and observed
gpuProfile; state which reads ran. - Route to scheduling, observability, KAITO, or scaling.
- Separate container/host OOM from device-allocation failures. For
OOMKilled, correlate container memory limits/usage and node conditions. For Managed+Install device-memory evidence, use the exporter on port 19400 and incident-windowFB_USED/FB_FREE; sizing tables do not establish a cause. - For KAITO not-ready, warn that Workspace deletion leaves its GPU pools; cleanup is separate and needs explicit authorization.
- Report evidence, confidence, missing evidence, and owner handoff.
Require explicit authorization before scaling, cordon/drain, deletion, add-on enablement, or monitoring mutation.
Error Handling
| Condition | Response |
|---|---|
| Missing/contradictory evidence | Mark unconfirmed; request the owner/read |
| Unknown profile | Preserve evidence; do not prescribe stack repair |
| Mutation required | Propose separately and wait for authorization |
权限
docs.pytorch.orglearn.microsoft.com检查
低风险 · 没有发现需要提醒的地方。
未经人工审核 · 已做规则检查;模型审核尚未开启。
文件5 个文件 · 12.2 KB
- SKILL.md3.0 KB
references/4
- gpu-cost-and-scaling.md1.6 KB
- gpu-observability.md2.7 KB
- gpu-scheduling.md3.2 KB
- kaito-workspaces.md1.8 KB
版本
- #11.1.1最新2026年10月10日