Brand promo video generator

01What is it?
For marketers and creators producing promotional content for brands, products, websites, apps, shops, or personal projects. Its edge is a particular angle on brand and messaging, giving the agent tighter constraints than a plain brand promo video generator request.
02Inputs
Context for brand and messaging: your goals, audience, constraints, and any source material the skill asks for.
03Output
A ready-to-use result for brand and messaging: the analysis, copy, or recommendations the agent produces.
Install-only

Install as a package

Installs this one skill package for your coding agent, including any supporting files that skill ships with — not every skill in the repository. Read the tutorial.

Terminal
$ npx skills add minimax-ai/minimax-h3 --skill brand-promo-video-generator

Skill instructions

The instruction file for this skill. The skill also includes other files you need to install to use it.

SKILL.md

Brand Promo Video Generator

Create a polished short promo video for a brand, product, website, app, shop, or personal project. Use this Skill when the user has a logo, product images, screenshots, a website link, or just a clear idea and wants the agent to turn those materials into a clean brand reel.

This Hub adaptation replaces third-party Vibe Motion / Remotion implementation details with Hub-native orchestration: source research, asset verification, story planning, image/video generation, optional speech or music, and editing assembly. Do not initialize external projects or depend on npm validators during normal execution.

Tool Coverage Rule

The allowed-tools list must cover the full production promise in this Skill: source lookup, image preparation/generation, video clip generation, optional speech/music/audio generation, final editing/assembly, media inspection, and canvas grouping. If a runtime does not provide one of the listed generation or editing tools, downgrade the deliverable explicitly to a pre-production package instead of claiming that a final promo video can be generated.

STEP 1: Intake assets and resolve the brief

Before any story planning or generation, run a required user intake. Ask the user to upload or provide links to the elements that must be verified:

  • Logo files or official logo source pages
  • Font files, font names, typography guidance, or official website pages that show brand typography
  • Brand colors, color system, style guide, or pages that clearly show official colors
  • Product images, UI screenshots, packaging, renders, footage, or other brand imagery
  • Product information: official product name, feature list, launch focus, claims, CTA, disclaimers, and target audience
  • Company/product official URL or official source package

Treat user-uploaded assets as usable by default for pre-production and concept planning. Do not repeatedly ask a standalone rights/permission question during the first intake unless there is a concrete risk signal, such as a visible third-party watermark, obviously scraped marketplace imagery, contradictory user wording, legal/medical/financial compliance claims, or the user asks for commercial publication. Record the source as "user-provided" and surface rights caveats in the source summary instead of blocking the flow.

In the same opening intake, ask the user to choose:

  • Target duration, normally 15-30 seconds; recommend 15 seconds when the user wants a fast launch film
  • Aspect ratio; offer common choices such as 16:9, 9:16, 1:1, 4:3, 3:4, or match a supplied reference

Also identify campaign focus, distribution channel, narration language, on-screen copy language, and visible copy needs when they are not already clear. Do not proceed to creative direction until the user has supplied the usable materials or explicitly confirms which elements are unavailable.

Language rule for promo content: choose narration and on-screen copy language from the brand materials, target audience, and platform context, not mechanically from the chat language. If the brand assets and visible source copy are primarily English or global corporate English, default narration and on-screen copy to English unless the user explicitly asks for Chinese localization. If the user is Chinese but says they are testing as a new user, keep chat replies in Chinese, but plan the actual video copy in the language that best fits the brand campaign.

If a logo, product UI, person, mascot, packaging, font, color system, or other identity-bearing asset cannot be authenticated, stop and ask for an authorized original instead of generating a plausible substitute.

STEP 2: Build the brand truth sheet

Research or inspect the strongest available sources and summarize the brand truth sheet before creative production:

  1. User-provided original exports
  2. Official company website, static bundle, newsroom, brand portal, media kit, press kit, or official repository
  3. Company-controlled media library
  4. Licensed stock or authorized partner kit

Extract:

  • Exact logo variants, clear space, and usage constraints
  • Official fonts or visible typographic behavior
  • Primary, secondary, and dynamic brand colors
  • Brand tone, principles, visual motifs, and interaction language
  • Current product names, features, scenarios, metrics, slogans, CTA, and disclaimers
  • Official photography, renders, UI screenshots, footage, press assets, and media kit material

Do not use logo aggregation sites, search thumbnails, fan recreations, Pinterest reposts, or AI-generated substitutes as identity-bearing sources. Procedural graphics are allowed only as non-representational motion layers: masks, gradients, color fields, grids, glows, particles, trails, typography, verified-data charts, and transition geometry.

STEP 3: Create a provenance manifest

Record every identity-bearing asset in a compact manifest that can be delivered to the user. The manifest may be a text node or table and should include:

  • Stable asset ID and role
  • Local path or canvas node reference when available
  • Exact source URL or user-provided source note
  • Source type, such as official website, media kit, user-provided original, or licensed stock
  • Verification target used for comparison
  • Rights or publication note
  • Authenticity status: verified, user-supplied, licensed, or blocked

The manifest does not grant publication rights by itself. If authorization is unclear, label the video as an unofficial concept and tell the user commercial publication requires permission.

STEP 4: Choose the story spine

Present 2-3 concise creative directions when the user has not already chosen one, recommend one, and continue after confirmation. Use the product category to pick a spine:

  • AI / SaaS: user intent -> thinking or planning -> capabilities -> execution -> useful output -> proof -> logo
  • Physical product: hero reveal -> interaction -> feature macro -> usage context -> result -> logo
  • Service / company: context -> process -> evidence -> outcome -> promise -> logo
  • Image-led brand: authentic imagery -> visual motif -> benefit -> emotional payoff -> logo

Keep the story product-specific. Show actual features, interactions, scenarios, outputs, and proof instead of hiding the story behind abstract effects.

STEP 5: Plan exact beats

Plan a frame-aware timeline before generation. Use 30fps as the planning convention unless the output pipeline requires otherwise.

For a 15-second film, target 5-8 major beats. For a 30-second film, target 8-12 major beats. Each beat should define:

  • Start and end time or frame range
  • Visual owner and authentic asset IDs
  • Primary action
  • Product or brand proof shown in the shot
  • Copy and readable hold
  • Color state
  • Incoming and outgoing transition
  • Motion intent: setup, anticipation, commitment, impact, brake, settle

A useful 15-second pattern is: brand hook, user intent or setup, product mechanism, capabilities or scenarios, output or proof, product payoff, final logo and CTA. Use 6-12 frame overlaps when outgoing motion naturally supplies the next shot.

STEP 6: Direct the motion language

Build intensity with control:

  • Let product motion, cursor paths, UI flow, light, scrolling content, object edges, or matched geometry drive transitions
  • Use 2-5 deliberate color states tied to meaning
  • Keep one primary action per beat; delay secondary layers slightly
  • Establish 2-3 high-energy peaks and quieter braking moments
  • Preserve readable silhouettes, copy, and logo clear space
  • Avoid fake HUDs, arbitrary glass cards, decorative text walls, unverified metrics, and identical easing everywhere

For AI products, include at least one readable chain such as: prompt -> planning -> parallel capabilities -> generated result -> proof. For physical or service products, show cause and effect from user action to concrete outcome.

STEP 7: Hard confirmation before generation

Before generating any video, image sequence, speech, music, or final edit, stop and show the user the completed pre-production package:

  • Provenance manifest or source summary
  • Brand truth sheet
  • Chosen creative direction
  • Exact beat / shot plan
  • Visible copy, CTA, narration, and audio plan
  • Known authenticity, rights, or placeholder caveats

Use a concise confirmation step before generation. If the user clearly expresses approval or intent to proceed after seeing the pre-production package — for example "confirm", "generate", "go ahead", "continue", "next", "可以", "继续", "下一步", or similar — treat it as permission to generate, unless the message also asks for changes. If the user asks to skip the process, still provide a compact source summary, brand truth sheet, and beat plan first, then proceed when they indicate approval. Offer revision choices only when the user's reply is ambiguous or requests changes.

STEP 8: Produce Hub assets

Use Hub-native generation and editing only after the hard confirmation gate has passed. Keep each dispatch self-contained with the chosen model, aspect ratio, authentic reference paths, and original request.

Typical production flow:

  1. Generate or prepare verified still frames, UI plates, product hero frames, or motion-ready story images.
  2. Generate video clips from those frames or from precise text prompts, preserving the same ratio and brand assets.
  3. Default audio policy for native brand reels: when the user asks for BGM, music, soundtrack, ambient sound, or says nothing beyond needing a finished promo video, prefer video-native audio from the selected video model instead of generating a separate music track. For the default MiniMax H3 route, set native audio on (generate_audio=true) and prompt for brand-safe instrumental music / UI sound design inside the video prompt.
  4. Generate separate speech or music only when the user explicitly needs controllable narration, voiceover, dialogue, replaceable standalone BGM, exact music duration independent of the video, post-production remixing, or when the selected video model cannot generate suitable audio. Do not duplicate a soundtrack between video-native audio and separate audio generation.
  5. Assemble clips, any explicitly separate audio, and final brand lockup in editing. Add subtitles only when the user explicitly asks for subtitles.

Do not redraw or approximate logos, wordmarks, product UI, packaging, mascot, person, or brand scene. Use generated material only for abstract motion, atmosphere, transition geometry, or clearly conceptual scenes that do not impersonate official product evidence.

STEP 9: Verify before delivery

Before final response, check:

  • The logo and identity-bearing assets came from verified or user-authorized sources
  • Product names, feature wording, claims, metrics, slogans, and CTA match official sources or are clearly marked as concept copy
  • The video duration, aspect ratio, and language match the brief
  • Copy is readable and not overcrowded
  • The final logo is not stretched, cropped, or rebuilt
  • Motion has clear visual ownership and does not obscure the product
  • The output is on the canvas and multi-asset outputs are grouped

If the output fails an authenticity check, replace the questionable asset with an official/user-authorized source or stop and ask for the asset. Never improve an imitation.

STEP 10: Deliver

Provide:

  • Final video path or canvas output
  • Duration, aspect ratio, and language
  • Short creative summary
  • Provenance manifest or source summary
  • Rights/disclaimer note when needed
  • Specific suggestions for the next iteration, such as pacing, claim clarity, CTA, audio, or platform crop

Failure recovery

  • Wrong or approximate logo: remove it, locate the current official file or ask for the user's original, then regenerate or re-edit.
  • Fake-looking product/UI: replace with official, user-supplied, or licensed media. Do not polish the imitation.
  • Beautiful but generic: add a complete product interaction, verified claim, or real output.
  • Fast but chaotic: reduce simultaneous actions, assign a visual owner, and preserve matched motion across cuts.
  • Smooth but slow: shorten holds, overlap transitions, and brake only around key messages.
  • Asset unavailable: ask for an authorized original; never guess.

Supporting file: meta.yaml

display-name-zh: 品牌宣传短片生成器
version: 0.1.9
tag-en: Commercial Ad
tag-cn: 商业广告
complete-tags-en:
- Commercial Ad / Planning
- Commercial Ad / Creative Generation
- Commercial Ad / Post-production
complete-tags-cn:
- 商业广告 / 计划制定
- 商业广告 / 创作生成
- 商业广告 / 后期处理
summary-en: Turn verified brand assets and campaign goals into a polished promotional short video.
summary-cn: 基于品牌素材与推广目标,完成事实核验、创意方向、分镜和音画合成,输出可展示的品牌宣传短片。
desc-en: For marketers and creators producing promotional content for brands, products, websites, apps, shops, or personal projects. Users provide logos, product images, interface screenshots, official links, or other verifiable assets and confirm duration, aspect ratio, audience, and campaign focus. The Skill organizes brand facts and asset provenance, selects a narrative direction, plans precise beats and shots, generates needed imagery, video, voiceover, or music, and completes assembly and pre-delivery review. It outputs a promotional short that highlights product capabilities, use cases, and a call to action. Best for launches, website showcases, and social promotion; not for imitating real brand marks without authorized assets, inventing product claims, or producing long-form narrative films.
desc-cn: 面向需要为品牌、产品、网站、App、小店或个人项目制作宣传内容的运营与创作者。用户需提供 LOGO、产品图、界面截图、官网链接或其他可核验素材,并确认时长、画幅、受众和推广重点。Skill 会整理品牌事实与素材来源,选择叙事方向,规划精确节拍和镜头,生成所需图像、视频、旁白或音乐,并完成合成与交付前检查。最终输出一条突出产品功能、使用场景和行动号召的品牌宣传短片。适用于新品发布、官网展示和社交媒体推广,不适用于缺少授权素材时仿造真实品牌标识、虚构产品功能或制作长篇剧情影片。
author-en: MiniMax Hub
author-cn: MiniMax Hub
source: official-featured

Supporting file: SKILL.cn.md

品牌宣传短片生成器

为品牌、产品、网站、App、小店或个人项目制作一条好看、清晰、能直接展示的宣传短片。用户可以提供 LOGO、产品图、截图、官网链接,也可以只给一个想法;Skill 会把这些素材整理成品牌短片。

此 Hub 适配版已移除第三方 Vibe Motion / Remotion 项目初始化与 npm 校验依赖,改为 Hub 原生编排:来源研究、资产核验、故事规划、图像/视频生成、可选语音或音乐,以及后期合成。正常执行时不要初始化外部项目。

步骤 1:上传素材并确认简报

在任何故事规划或生成之前,必须先进行用户问询。请用户上传或提供链接,用于核验以下元素:

  • LOGO 文件或官方 LOGO 来源页面
  • 字体文件、字体名称、字体规范,或能体现品牌字体的官方网站页面
  • 品牌颜色、色彩系统、风格指南,或能清楚体现官方色彩的页面
  • 产品图片、界面截图、包装、渲染图、视频素材或其他品牌图片
  • 产品信息:官方产品名称、功能列表、发布重点、宣传主张、行动号召、免责声明和目标受众
  • 公司/产品官网或官方资料包

默认把用户上传的素材视为可用于前期方案和概念制作,不要在第一轮问询中反复单独追问“是否有授权”。只有出现明确风险信号时才询问,例如可见第三方水印、明显抓取的电商/图库图、用户表述互相矛盾、涉及法律/医疗/金融合规主张,或用户明确要求商用发布。来源清单中将素材记录为“用户提供”,并在来源摘要里提示必要的权利注意事项,而不是阻塞流程。

在同一次开场问询中,让用户选择:

  • 目标时长,通常为 15-30 秒;如果用户想要快速发布片,推荐 15 秒
  • 画幅比例;提供 16:9、9:16、1:1、4:3、3:4 或匹配用户参考图等常见选项

如果活动重点、投放渠道、旁白语言、画面文案语言和画面文案需求尚不清楚,也在此阶段一并确认。用户未提供可用素材,且未明确说明哪些元素不可用之前,不要进入创意方向阶段。

宣传片内容语言规则:旁白和画面文案语言应根据品牌素材、目标受众和投放平台判断,不要机械跟随聊天语言。如果品牌素材和可见来源文案主要是英文或全球企业英文,默认旁白和画面文案使用英文,除非用户明确要求中文本地化。即使用户用中文聊天或说“我是新用户测试”,聊天回复仍用中文,但视频实际文案应使用最适合该品牌传播的语言。

如果 LOGO、产品界面、人物、吉祥物、包装、字体、色彩系统或其他带身份识别的资产无法认证,停止并向用户索要授权原件,不要生成看起来相似的替代品。

步骤 2:建立品牌事实表

研究或检查最可靠的来源,并在创意制作前总结品牌事实表:

  1. 用户提供的原始导出文件
  2. 官方网站、静态资源包、新闻中心、品牌门户、媒体资料包、新闻资料包或官方仓库
  3. 公司控制的官方媒体库
  4. 授权图库或授权合作方资料包

提取:

  • 精确 LOGO 版本、安全空间和使用限制
  • 官方字体或可见的字体行为
  • 主色、辅助色和动态色彩
  • 品牌语气、原则、视觉母题和交互语言
  • 当前产品名称、功能、场景、指标、口号、行动号召和免责声明
  • 官方摄影、渲染图、界面截图、视频素材、新闻素材和媒体包内容

不要使用 LOGO 聚合站、搜索缩略图、粉丝重绘、Pinterest 转发或 AI 生成替代品作为带身份识别的来源。程序化图形只可作为非写实动效层:遮罩、渐变、色块、网格、辉光、粒子、轨迹、字体、已验证数据图表和转场几何。

步骤 3:创建来源清单

为每个带身份识别的资产记录一份简洁来源清单,可作为文字节点或表格交付给用户。清单应包含:

  • 稳定资产 ID 与角色
  • 本地路径或画布节点引用
  • 精确来源 URL,或用户提供来源说明
  • 来源类型,例如官网、媒体包、用户原始素材或授权图库
  • 用于比对的验证目标
  • 权利或发布说明
  • 真实性状态:已验证、用户提供、已授权或阻塞

来源清单本身不等于发布授权。授权不明确时,将视频标注为非官方概念,并告知用户商业发布需要许可。

步骤 4:选择故事脊柱

如果用户尚未选定方向,给出 2-3 个简短创意方向,推荐一个,并在确认后继续。按产品类型选择叙事脊柱:

  • AI / SaaS:用户意图 -> 思考或规划 -> 能力 -> 执行 -> 有用输出 -> 证明 -> LOGO
  • 实体产品:英雄展示 -> 交互 -> 功能特写 -> 使用场景 -> 结果 -> LOGO
  • 服务 / 公司:背景 -> 流程 -> 证据 -> 成果 -> 承诺 -> LOGO
  • 图像主导品牌:真实影像 -> 视觉母题 -> 利益点 -> 情绪回报 -> LOGO

故事必须产品专属。展示真实功能、交互、场景、输出和证明,不要用抽象特效掩盖薄弱产品叙事。

步骤 5:规划精确节拍

生成前先规划带帧意识的时间线。除非输出流程另有要求,默认用 30fps 作为规划单位。

15 秒影片建议 5-8 个主要节拍;30 秒影片建议 8-12 个主要节拍。每个节拍应定义:

  • 开始与结束时间或帧范围
  • 视觉主导元素和真实资产 ID
  • 主要动作
  • 镜头中展示的产品或品牌证明
  • 文案与可读停留
  • 色彩状态
  • 入场与出场转场
  • 动作意图:铺垫、预备、承诺、冲击、制动、稳定

一个实用的 15 秒结构是:品牌钩子、用户意图或场景建立、产品机制、能力或应用场景、输出或证明、产品回报、最终 LOGO 与行动号召。当前一个动作能自然带出下一个镜头时,可使用 6-12 帧重叠。

步骤 6:导演动效语言

用可控方式建立强度:

  • 让产品动作、光标路径、界面流、光线、滚动内容、物体边缘或匹配几何驱动转场
  • 使用 2-5 个有意义的色彩状态
  • 每个节拍保留一个主要动作;次级层稍作延迟
  • 设置 2-3 个高能峰值和较安静的制动时刻
  • 保持剪影、文案和 LOGO 安全空间清晰可读
  • 避免虚假 HUD、随意玻璃卡片、装饰性文字墙、未经验证的指标,以及全片同一种缓动

AI 产品至少包含一条可读链路,例如:提示词 -> 规划 -> 并行能力 -> 生成结果 -> 证明。实体或服务产品则展示从用户动作到具体成果的因果关系。

步骤 7:生成前硬确认

在生成任何视频、图片序列、语音、音乐或最终剪辑之前,必须停下,并先向用户展示完整前期方案:

  • 来源清单或来源摘要
  • 品牌事实表
  • 已选创意方向
  • 精确节拍 / 分镜计划
  • 画面文案、CTA、旁白和声音方案
  • 已知的真实性、授权或占位说明

生成前保持简洁确认。用户在看过前期方案后,如果明确表达认可或推进意图,例如“确认”“生成”“可以”“继续”“下一步”“go ahead”“continue”“next”等,应视为允许生成,除非同一句话同时要求修改。若用户要求跳过流程,仍需先给出简版来源摘要、品牌事实表和分镜节拍;随后用户表达同意推进即可生成。只有当用户回复含糊或提出修改时,才列出修改分镜、修改文案、修改声音或暂停等选项。

步骤 8:生产 Hub 资产

只有生成前硬确认通过后,才可以使用 Hub 原生生成与剪辑。每次派发都应自包含:所选模型、画面比例、真实参考路径和用户原始需求。

典型制作流程:

  1. 生成或准备已验证的静帧、界面底板、产品主视觉帧或适合动效的视频首帧。
  2. 基于这些画面或精确文字描述生成视频片段,并保持相同比例与品牌资产。
  3. 品牌短片的默认声音策略:当用户要求 BGM、音乐、配乐、氛围声,或只说要一条完整宣传片而没有要求独立音轨时,优先使用所选视频模型的原生音频,不要默认拆成“无声视频 + 单独音乐”。默认 MiniMax H3 路线应开启原生音频(generate_audio=true),并在视频提示词中写入品牌安全的纯音乐 / UI 音效设计。
  4. 只有当用户明确需要可控旁白、口播、对白、可单独替换的 BGM、独立于视频的精确音乐时长、后期混音,或所选视频模型无法生成合适音频时,才单独生成语音或音乐。不要让视频原生音频和独立音频重复生成同一条声轨。
  5. 在剪辑阶段合成片段、用户明确要求的独立音频和最终品牌锁定画面。只有用户明确要求字幕时才添加字幕。

不得重绘或近似 LOGO、字标、产品界面、包装、吉祥物、人物或品牌场景。生成素材只用于抽象动效、氛围、转场几何,或明确为概念化且不冒充官方产品证据的场景。

步骤 9:交付前验证

最终回复前检查:

  • LOGO 和带身份识别的资产来自已验证或用户授权的来源
  • 产品名称、功能表述、主张、指标、口号和行动号召与官方来源一致,或被明确标注为概念文案
  • 视频时长、比例和语言符合简报
  • 文案可读且不过度拥挤
  • 最终 LOGO 未被拉伸、裁切或重建
  • 动效有清晰视觉主导,不遮挡产品
  • 输出已在画布上,多资产输出已编组

如果真实性检查失败,用官方或用户授权来源替换可疑资产,或停止并向用户索取资产。绝不优化仿制品。

步骤 10:交付

提供:

  • 最终视频路径或画布输出
  • 时长、比例和语言
  • 简短创意总结
  • 来源清单或来源摘要
  • 必要的权利/免责声明
  • 下一轮可改进的具体建议,例如节奏、主张清晰度、行动号召、音频或平台裁切

失败恢复

  • LOGO 错误或近似:移除它,定位当前官方文件或向用户索要原件,再重新生成或剪辑。
  • 产品/界面看起来虚假:替换为官方、用户提供或授权素材,不要美化仿制品。
  • 好看但通用:加入完整产品交互、已验证主张或真实输出。
  • 快但混乱:减少同时动作,指定视觉主导,并保持跨镜头匹配运动。
  • 顺滑但太慢:缩短停留,重叠转场,只在关键信息处制动。
  • 资产不可用:索要授权原件,绝不猜测。

How do I install Brand promo video generator in Cursor, Claude Code, or Codex?

Run npx skills add minimax-ai/minimax-h3 --skill brand-promo-video-generator in the project where you want it, then ask your agent for the skill by name. The --skill flag installs only Brand promo video generator, not every skill in the repository.

Where does Brand promo video generator come from and what license is it under?

Brand promo video generator comes from the minimax-ai/minimax-h3 repository on GitHub. That repository has 7.1K GitHub stars. No license was detected on the source repository, so check with the author before redistributing it.

Prefer plain text? Read the Brand promo video generator guide as markdown.