Token导航 LogoToken导航TokenDH.com
研究检索敏感数据github未标认证来源可访问clear审计未展示

nano-banana-video-generation纳米香蕉视频生成

Agent Skill

用于辅助视频生成、动画合成、脚本化剪辑或 Remotion 等视频项目开发。它适合让 Agent 组织镜头、生成素材说明、维护合成代码或排查渲染问题。使用时需要确认分辨率、时长、素材路径和导出格式;涉及外部素材、人物肖像或商业发布时,应先核对版权授权和内容审核要求。

总安装

13,293

周安装

316

GitHub Stars

公开资料未说明

下载量

4,475
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:nano-banana-video-generation(纳米香蕉视频生成)
来源仓库:https://github.com/the-focus-ai/nano-banana-cli
仓库路径:skills/nano-banana-video-generation
安装命令:
npx skills add the-focus-ai/nano-banana-cli --skill "nano-banana-video-generation"
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

AgentSkills.tonpx skills
npx skills add the-focus-ai/nano-banana-cli --skill "nano-banana-video-generation"

简介

用于辅助视频生成和动画合成项目开发。nano-banana-video-generation 属于研究检索类 Skill,可作为该场景下的辅助能力补充。

  • 适合让 Agent 组织镜头、维护合成代码或排查渲染问题。
  • 使用时需确认分辨率、时长和导出格式等参数。
  • 涉及人物肖像或商业发布时应先核对版权和内容审核要求。
  • 安装方式:通过 GitHub 仓库添加,支持 Codex、Claude 等宿主。

SKILL.md

Nano Banana Video Generation

Generate videos using Google Veo 3.1 models via the nano-banana CLI.

Prerequisites

  • GEMINI_API_KEY environment variable must be set
  • The CLI is installed via npx @the-focus-ai/nano-banana

Quick Reference

# Generate a video from text
nano-banana --video "A sunset over mountains, slow dolly-in, cinematic lighting"

# Animate an existing image
nano-banana --video "The character slowly turns and smiles" --file portrait.png

# Cost-optimized development mode
nano-banana --video "Quick test scene" --video-fast --no-audio --resolution 720p

# Specify output path
nano-banana --video "A cat playing" --output cat-video.mp4

# Full control over settings
nano-banana --video "Dramatic reveal scene" \
  --duration 8 --aspect 16:9 --resolution 1080p --seed 42

Understanding Video Requests

Before generating, clarify these video-specific aspects:

  1. Core Scene: What's the main action or subject?
  2. Camera Movement: Static, dolly, pan, tracking, crane?
  3. Style: Cinematic, documentary, commercial, casual?
  4. Audio: Dialogue? Sound effects? Ambient sounds? Music?
  5. Duration: 4, 6, or 8 seconds?
  6. Orientation: Landscape (16:9) or portrait (9:16)?

The Five-Part Video Prompt Formula

Structure prompts with these elements:

[Camera Movement] + [Subject] + [Action] + [Environment] + [Audio/Style]

Example - Weak prompt:

"a person walking"

Example - Strong prompt:

"Slow dolly-in shot. A woman in her 30s, shoulder-length wavy black hair,
green jacket, walks confidently through a sunlit park. Golden hour lighting,
warm color grading. Ambient sounds: birds chirping, distant traffic.
Cinematic, aspirational mood. No subtitles, no text overlay."

Workflow

Step 1: Craft the Prompt

Use the prompting-guide.md for comprehensive guidance.

Key principles:

  1. Start with camera movement (dolly, pan, static, tracking)
  2. Describe subject in detail (appearance, wardrobe, expression)
  3. Specify action with timing cues
  4. Include lighting and environment
  5. Add audio design (dialogue, SFX, ambient)
  6. Always end with: "No subtitles, no text overlay, no captions"

Step 2: Consider Cost

Video generation is significantly more expensive than images:

ModelCost per Second8-Second Video
veo-3.1-generate-preview$0.50-0.75$4-6
veo-3.1-fast-generate-preview$0.10-0.15$0.80-1.20

Development workflow:

  1. Iterate with --video-fast --no-audio (cheapest)
  2. Test with --video-fast (add audio when needed)
  3. Final render with default model (premium quality)

Step 3: Generate

nano-banana --video "your detailed prompt here"

Generation takes 2-4 minutes. Progress is shown in the terminal.

Step 4: Iterate

If the result isn't right:

  1. Refine camera movement - Be more explicit (e.g., "slow dolly-in over 8 seconds")
  2. Add negative guidance - Describe what to avoid
  3. Simplify - Focus on one main action per clip
  4. Try different duration - 4s or 6s may work better for quick actions

Commands

Text-to-Video

nano-banana --video "<prompt>"

Image-to-Video (Animation)

nano-banana --video "<motion description>" --file <input-image>

The motion description should describe how the image should animate:

  • "The character slowly turns their head and smiles"
  • "The scene comes alive with subtle wind movement"
  • "Zoom out to reveal the full landscape"

Options

OptionDescriptionDefault
--videoEnable video mode(required)
--video-model <name>Veo model to useveo-3.1-generate-preview
--video-fastUse fast/cheap model(premium model)
--duration <sec>4, 6, or 8 seconds8
--aspect <ratio>16:9 or 9:1616:9
--resolution <res>720p or 1080p1080p
--audioGenerate audio(enabled)
--no-audioDisable audio-
--seed <number>Reproducibility seed(random)
--output <file>Output pathoutput/video-.mp4
--file <image>Input image to animate-

Camera Movement Reference

Use these terms for precise camera control:

MovementDescriptionExample Prompt
StaticNo movement"Static shot on tripod. A coffee cup steaming..."
PanHorizontal rotation"Slow pan left across the city skyline..."
TiltVertical rotation"Tilt down from face to hands..."
Dolly InCamera moves closer"Slow dolly-in from medium to close-up..."
Dolly OutCamera moves away"Dolly-out revealing the vast landscape..."
TrackingParallel to subject"Tracking shot following character walking..."
CraneSweeping vertical"Crane shot ascending from ground level..."
HandheldRealistic shake"Handheld camera, documentary style..."

Important: Use ONE primary movement per shot. Don't combine multiple movements.

Dialogue Formatting

For spoken dialogue, use the colon format:

Character description says: "Exact dialogue here."

Example:

"A friendly young woman, excited and cheerful, says: 'Welcome to our store!'
Standing in bright retail environment. Natural lip-sync. No subtitles."

Guidelines:

  • Keep dialogue to 6-12 words for 8 seconds
  • Describe the speaker's tone and emotion
  • Always add "No subtitles, no text overlay"

Audio Design

Structure audio in layers:

  1. Dialogue (highest priority) - Always clear
  2. Sound Effects - Specific, timed actions
  3. Ambient - 3-5 background elements max
  4. Music - Lowest priority, "ducks under dialogue"

Example:

"Sound effects: Door closing at 2-second mark, footsteps on wood.
Ambient sounds: Quiet office hum, distant typing.
Background music: Soft jazz, low volume, ducks under dialogue."

Best Practices

For Better Results

  1. Front-load important info - Camera, subject, action first
  2. Use cinematic terms - "35mm lens", "shallow depth of field", "golden hour"
  3. Be specific about lighting - "Soft window light from left", not just "good lighting"
  4. Describe the mood - "Intimate", "epic", "suspenseful", "uplifting"
  5. Include negative guidance - What to avoid

For Image-to-Video

  1. Match the image - Describe motion that fits what's in the image
  2. Start subtle - Small movements work better than dramatic changes
  3. Keep lighting consistent - Don't describe lighting changes that differ from the image

For Consistency Across Shots

When creating multiple related videos:

  1. Create a character description and reuse it exactly
  2. Keep lighting style consistent
  3. Use the same camera movement style family
  4. Use --seed for more reproducible results

Troubleshooting

"Video generation timeout"

  • Generation can take 2-4 minutes
  • If persistent, try simpler prompts
  • Use --video-fast for faster generation

Poor quality or wrong content

  • Add more specific descriptions
  • Include negative guidance
  • Try the premium model instead of fast

Subtitles appearing in video

  • Always include "No subtitles, no text overlay, no captions" in prompt
  • Veo was trained on videos with subtitles and tends to add them

Audio doesn't match video

  • Be more specific about when sounds occur
  • Use "Sound effect: X at Y-second mark"
  • Simplify audio layers (fewer elements)

Safety filter rejection

  • Avoid violence, weapons, explicit content
  • Rephrase ambiguous terms
  • Try more generic descriptions

Cost Optimization

# Development (cheapest): ~$0.80 per video
nano-banana --video "test prompt" --video-fast --no-audio --resolution 720p

# Testing with audio: ~$1.20 per video
nano-banana --video "test prompt" --video-fast

# Production quality: ~$6 per video
nano-banana --video "final prompt" --resolution 1080p

Example Prompts

See the examples/ directory for complete prompt examples:

Environment Setup

Ensure GEMINI_API_KEY is set:

export GEMINI_API_KEY="your-api-key-here"

Or create a .env file in your project:

GEMINI_API_KEY=your-api-key-here

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

04

需要参考平台分布和安装热度时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Cursor

74.89%
按下载量换算3,351

安全审计

暂无安全审计结果可展示。

权限和风险

敏感数据

该 Skill 可能接触密钥、Token、环境变量或敏感配置,应进入高风险复核队列,默认不自动发布。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills