Token导航 LogoToken导航TokenDH.com
运维和基础设施需要联网github未标认证来源可访问许可证需确认审计异常

nebius-dedicated-endpointnebius 专用端点

Agent Skill

用于辅助 API 设计、接口文档、请求响应结构和服务集成说明。它适合让 Agent 梳理 endpoint、生成 OpenAPI 草稿、检查字段命名、整理错误码或辅助前后端联调。使用时需要确认真实业务语义、鉴权方式、分页和错误处理规则;涉及生成接口文档时,应避免凭空补字段,最好从现有代码、schema 或接口样例中提取事实。

总安装

196

周安装

8

GitHub Stars

3

下载量

63
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:nebius-dedicated-endpoint(nebius 专用端点)
来源仓库:https://github.com/arindam200/nebius-skills
仓库路径:skills/nebius-dedicated-endpoint
安装命令:
npx skills add https://github.com/arindam200/nebius-skills --skill nebius-dedicated-endpoint
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/arindam200/nebius-skills --skill nebius-dedicated-endpoint

简介

nebius-dedicated-endpoint 用于辅助 API 设计、接口文档和请求响应结构说明,适合梳理 endpoint 和生成 OpenAPI 草稿。

  • 适用于前后端联调、服务集成或接口规范制定等开发协作场景。
  • 通过 GitHub 仓库安装,使用 npx skills add 命令添加技能。
  • 涉及接口文档时应避免凭空补字段,建议从现有代码或样例中提取事实。
  • 适用宿主包括 Codex、Claude、Cursor、Gemini CLI,接入前应确认版本、权限和运行环境要求。

SKILL.md

Nebius Dedicated Endpoints

Dedicated endpoints give you an isolated, GPU-backed deployment of a supported model template with per-region data residency, configurable autoscaling, and OpenAI-compatible inference.

Prerequisites

pip install requests openai
export NEBIUS_API_KEY="your-key"

Control plane (manage endpoints): https://api.tokenfactory.nebius.com Data plane (inference), pick by region:

RegionInference base URL
eu-north1https://api.tokenfactory.nebius.com/v1/
eu-west1https://api.tokenfactory.eu-west1.nebius.com/v1/
us-central1https://api.tokenfactory.us-central1.nebius.com/v1/

Key concepts

  • Template — deployable blueprint (model + supported GPU types/regions)
  • Flavorbase (throughput-optimized) or fast (low-latency, speculative decoding)
  • Endpoint — your live deployment, identified by endpoint_id
  • routing_key — the model name to pass in inference calls

Operations

List available templates

import requests
r = requests.get("https://api.tokenfactory.nebius.com/v0/dedicated_endpoints/templates",
                 headers={"Authorization": f"Bearer {API_KEY}"})
templates = r.json().get("templates", [])
for t in templates:
    print(t["template_name"], [f["flavor_name"] for f in t.get("flavors", [])])

Create an endpoint

payload = {
    "name":     "my-endpoint",
    "template": "openai/gpt-oss-20b",      # from list_templates
    "flavor":   "base",
    "region":   "eu-north1",
    "scaling":  {"min_replicas": 1, "max_replicas": 2},
}
r = requests.post("https://api.tokenfactory.nebius.com/v0/dedicated_endpoints",
                  headers=HEADERS, json=payload)
endpoint = r.json()
endpoint_id  = endpoint["endpoint_id"]
routing_key  = endpoint["routing_key"]

Poll GET /v0/dedicated_endpoints/{endpoint_id} until status == "ready".

Run inference

from openai import OpenAI
client = OpenAI(base_url="https://api.tokenfactory.nebius.com/v1/", api_key=API_KEY)

resp = client.chat.completions.create(
    model=routing_key,          # the routing_key from endpoint creation
    messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)

Update autoscaling (live, no downtime)

requests.patch(
    f"https://api.tokenfactory.nebius.com/v0/dedicated_endpoints/{endpoint_id}",
    headers=HEADERS,
    json={"scaling": {"min_replicas": 2, "max_replicas": 8}},
)

Delete endpoint

requests.delete(
    f"https://api.tokenfactory.nebius.com/v0/dedicated_endpoints/{endpoint_id}",
    headers=HEADERS,
)

Choosing flavor

NeedUse
High throughput, cost-efficientbase
Low latency, real-time UXfast (uses speculative decoding + smaller batches)

Data residency

Choose region to control where inference runs. Metrics are collected locally but stored in eu-north1.

Bundled reference

Read references/templates-regions.md when the user asks about available templates, GPU types, regions, or flavor differences.

Reference script

Full working script: scripts/02_dedicated_endpoints.py

Docs: https://docs.tokenfactory.nebius.com/ai-models-inference/dedicated-endpoints

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

34.69%
按下载量换算22

Claude

28.44%
按下载量换算18

Cursor

20.02%
按下载量换算13

Gemini CLI

9.22%
按下载量换算6

安全审计

Gen Agent Trust Hub

通过

Socket

未通过

Snyk

通过

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills