返回文章

Seedance 2.5 AI 生成影片完整教學

作者:Frankie Chan(Kick Ads 聯合創辦人) · 更新於 2026年8月13日

Seedance 2.5 AI 生成影片完整教學

同一日、同一個帳戶,我們用 Seedance 2.5 生成了兩條片。第一條 prompt 寫明要 cinematic commercial 質感的菠蘿油,出來像個塑膠模型,第二條寫茶餐廳沖奶茶,枱面有茶漬深到變色的絲襪袋、燒到變色的鋼壺、防火膠板上一堆杯印和一灘水,出來就像間開了二十年的茶記。

左邊清空了場景,右邊把雜物放回枱面,同樣是 Seedance 2.5 生成

第二條 prompt 形容外觀的字眼反而比第一條少,差別在枱面上有多少東西。

簡答: Lumina 按秒收費,720p 每秒 46 credits。一條 20 秒廣告一鏡到尾生成收 920 credits,拆成三條 8 秒再剪起是 1,104,貴在多買了要剪走的片頭片尾。尾卡在 Canva 做,不另收生成費。

下面六個步驟由兩張參考相開始,到剪出一條可以直接上載的 20 秒廣告,文中有兩條完成品,也寫明每一步實際用了多少 credits,後半段另有十個行業的 prompt 可以直接複製。想知道 Seedance 2.5 本身值多少錢、生成文字為甚麼救不到、廣東話配音夠不夠出街,看Seedance 2.5 完整實測

你可以在香港用哪個 AI 工具生成影片?

你去搜尋 AI 生成影片,第一頁多數是剪片工具,例如 Canva、CapCut、Adobe,這些工具處理你已經有的片段,並不會憑文字畫出新畫面。要由零生成一段從來沒有拍過的影片,可以用的是下面這幾個。

工具香港可否直接付費生成一條 8 秒 720p適合
BytePlus Lumina(Seedance 2.5)可以,國際信用卡368 credits產品片、實景質感、廣東話配音
即夢 Dreamina內地版,香港付款有阻滯另一套計價,我們未測試內地市場
Sora、Veo可以各自計價,另計創意短片為主

本文的影片全部在 Lumina 生成,因為香港可以直接註冊和付款,而且它支援用參考圖鎖住產品。想比較 Seedance 2.5 本身的表現,看Seedance 2.5 香港實測

為甚麼有些 AI 生成影片一看就假?

Seedance 2.5 寫不到字,所以菠蘿油那條的 prompt 索性把場景清空,不准出現餐牌、價錢卡和海報,背景矇到看不見,結果就是一件食物放在一片啡色的散焦背景前面,畫面裡再沒有第二樣東西。那個包本身也錯了,菠蘿油應該橫切一刀把牛油打橫夾進去,這一條卻由中間直劈開,牛油豎起來插在中間,香港的茶記不會這樣出餐。

一條 AI 生成影片實際要多少錢

Lumina 按秒收費,720p 每秒 46 credits 而直片橫片價錢一樣。

設定Credits約港幣
720p・4 秒184HK$7
720p・8 秒(十條參考庫片用這個)368HK$14
720p・10 秒460HK$18
720p・20 秒(一鏡到尾)920HK$35
720p・30 秒(單鏡測試)1,380HK$53

我們用的是 Ultra plan,US$239 買 48,200 credits,一個 credit 折合約 HK$0.038,上表的港幣金額全部這樣計。

收費線性,所以一條 20 秒廣告的價錢只看你怎樣生成。一鏡到尾生成 20 秒是 920 credits,用三條 8 秒剪成是 1,104,貴 184,多出來的是每條片頭尾要剪走的那截。分開跑貴一點,但改動便宜,兩者的取捨在下面第二步講。

要分開跑就用 8 秒一條。一個做得完的動作要四至六秒,而 Seedance 2.5 頭半秒還在定位,尾半秒經常飄,剪走之後剩下六七秒乾淨畫面剛好夠用。10 秒貴四分之一,多出那截我們照樣剪走。

想自己核對價錢,先跑一條 4 秒的測試片,184 credits 就知道。

用 AI 生成影片做廣告的六個步驟

下面每一步都在自己付費的帳戶上跑過,文末兩條廣告就是這樣做出來的。

開始之前:開戶和買 credits

先去 Lumina 開個帳戶,用普通電郵就註冊到,國際信用卡付款,credits 即時到帳。我們用的是 Ultra plan,US$239 買 48,200 credits。

介面有 Creation 和 Director 兩種 mode,我們的影片全部用標準 Creation,文字出片、image-to-video 和掛參考圖都在這裡做。

Lumina 的模型選單,Seedance 2.5 之下有 Image To Video、Text To Video、Extend Video

比例、解像度和長度都在設定面板選。

Lumina 的 video settings 面板,比例、解像度與長度都在這裡選

按生成之前,按鈕上會先顯示這一 run 要用多少 credits(8 秒 720p 顯示 368)。

第一步:準備兩張參考相

一張影產品,即是你賣的東西,可以是菠蘿油、一支精華或者一部車,另一張影場地,即是那件東西平時放在哪裡,水吧枱、工場和貨架都可以。

兩張相不一定要重新影,用手機影就可以,現成的產品相一樣用得到,只要平光和夠悶。加了濾鏡或者景深很淺的相會把那套味道帶進之後每一條片,到時想改都改不到,所以修過的宣傳相反而不及一張隨手影的清楚。

產品那張正對鏡頭平影,label 向鏡,不要有東西遮住,場地那張則影闊一點,讓模型讀得到地面、牆身、光的方向,以及現場有多亂。

你給 Seedance 2.5 甚麼,它就還原甚麼,而參考圖的力度大過 prompt。我們上載過一張真車相,prompt 在五個位置寫明不要車徽、標誌、logo、字母和品牌記號,出來的片照樣有車廠徽章,保險桿上還多了一行自己作出來的字。

片裡出現真實品牌多數不是問題,汽車美容店的車有車廠徽章、茶餐廳枱面有支汽水,本身都很正常。麻煩的是另一種:你籠統形容一件產品,一個字都沒有提過牌子,Seedance 2.5 就自己造一個似模似樣的標誌出來,那個標誌不屬於任何人,卻明顯在模仿某個品牌。掛自己產品的真相讓它照著還原,或者寫明表面素身無標記,兩條路都避得開。

prompt 明文禁止過車徽,車頭的廠徽仍然清晰可見

第二步:將廣告分成三段,每段做一件事

開場、中段、收尾各一段,每段只負責一件事。

這三段可以寫成三條獨立 prompt 分開跑,也可以寫成一條 prompt 裡面三個階段,一鏡到尾生成。兩條路我們都試過。一鏡到尾那條收 1,380 credits,連戲完全過到,由第一格到最後一格都是同一部車、同一種光、同一塊漆,段與段之間的轉場由 Seedance 2.5 自己接,接得比我們手剪順。

分開三條的主要好處是改動便宜,一隻手畫錯,重跑一條蝕 368;一鏡到尾就要再花 920 重跑成條 20 秒。所以想同一個場景出幾個版本落廣告,就分開跑;prompt 寫得穩、一次收貨就出街,一鏡到尾反而更順。另外長度要寫得準,我們那條寫了 30 秒,但故事在第二十秒已經講完,最後十秒是白買的。

無論走哪一條,三段都要一次過寫完才開始生成,光位、時間和場景表面全部寫成一樣,因為連戲要在 prompt 裡面先定下來,後期補不回。

開場鏡頭交代場地 影闊一點,觀眾要看得出這是一間真的舖,產品在畫面裡面,第一格已經有東西在動,這個鏡頭的作用是令觀眾停低。

中段鏡頭是證據 鏡頭貼近手部動作,沖、抹、包都可以,觀眾真正會看完的就是這個鏡頭,我們也會在這裡多跑一次。

收尾鏡頭要停穩 產品放一邊,另一邊留白,結尾不要有快動作,因為價錢、offer 和 logo 之後要後期放進那片留白。

無論是三條一組還是單獨一條,每一條 prompt 都跟足四條規則。

一、不要清空場景 把雜物逐件寫出來,並且選不帶字的雜物,例如殘舊防火膠板、水印、剪落的花枝、濕抹布、崩口搪瓷、散落的碎料、地上的積水,而餐牌、價錢卡、包裝、白板和海報照樣要禁止,模型見到這些表面就會嘗試寫字。

二、手要入鏡,而且要在做事 前臂可以出現,人面就千萬不要寫,而觀眾就是靠手來判斷畫面真不真,所以 technical 那一段要寫明 hands have five fingers in every frame,每個 take 都要逐格檢查手指。

三、要有變化,而且變化要有後果 上面沖奶茶那條,茶柱倒落壺裡,泡就浮起來,八秒之內有一件事做完了。畫面維持一個完美狀態八秒,一看就知是 CGI,所以動作要寫成第一格已經在進行。

四、一條片只用一種運鏡 固定手持加輕微飄移、慢慢向下 tilt、橫向跟拍或者推近,選一種就夠。要避的是用慢慢推向主體來代替畫面裡真的有事發生,那是現成廣告素材的做法,一出現就多數看得出不是自己拍的。上面沖奶茶那條就是好例子,鏡頭定住不動,但茶柱一直在倒,泡一直在浮。下面那條 20 秒茶餐廳廣告中段用了推近,而牛油全程都沒有溶過,那一下推近就是推向一件靜止的東西,正正是這條規則要避的。

另外還有兩點,不要要求有錶盤、螢幕、摩打或者銘牌的機器,模型會自己發明一部出來,手工具亦只寫最簡單的,剪刀、水壺、洗車手套、刀。

第三步:生成時兩張參考圖要一併上載

三條 prompt 每一條都要掛住同一對參考圖,並且在 prompt 裡講明哪一張要照樣還原,哪一張用來做場景。

兩張圖都在 Image To Video 這個 mode 掛上去,按輸入框旁邊的 material 選單逐張加,一次加多過一張也可以。我們先加產品那張,再加場地那張,prompt 裡的 @image 1 和 @image 2 就跟這個次序寫,次序要自己記住,寫反了模型會把場地當成產品來還原。

Lumina 輸入框的 material 選單,用 local upload 逐張加自己的相

介面另外有一個 Main Body Reference 選項,我們測試期間一直按不動,兩張圖都是用普通附圖掛上去。

每條 prompt 開頭都先寫清楚分工:

@image 1 is the product: reproduce this exact item and its exact label
faithfully, do not redraw, restyle, re-space or re-colour it.
@image 2 is the location: reproduce this room, its surfaces, its light
direction and its layout, and stage the action inside it.

只寫文字不掛參考圖,三個鏡頭跑出來會是三間不同的舖、三部不同的車,剪在一起完全不像同一個地方。改成 image 1 綁產品、image 2 綁場地之後,三條分開跑的 8 秒片剪出來就是同一間舖。

輸出設定要在介面選,不要寫在 prompt 裡面,比例、解像度和長度先在面板選好,因為 prompt 寫著 9:16 而面板停在預設,生成出來會是橫片,credits 照樣扣。我們用 9:16、720p、8 秒。

prompt 裡面不要出現 --,後面的內容會被截斷,而截斷版照樣收足價錢。

被擋的 prompt 不扣 credit,內容過濾攔下來的時候,畫面只彈出「Try another prompt/image~」,沒有任何解釋,所以你要逐段貼回去試才找到元兇,好在不用花錢。很普通的字眼一樣會中招,我們寫 black mirror 就連續擋了五個版本,估計是撞上劇名。

第一次跑用 8 秒 720p 就夠,一條 368 credits。

第四步:先檢查成品,收貨的話才生成下一條

逐格看,對照下面七項,任何一項不過就重跑,不要嘗試補救。

檢查項這樣就作廢
主體在不在動作要對著的人或物根本不在畫面裡,手按著空氣或者一塊平布。
產品有沒有走樣與 image 1 對不上,比例變了、label 字距變了、多了本來沒有的細節。
顏色由參考圖帶進來的東西換了顏色,三條片駁在一起,這一項很容易穿崩。
任何一格不是五隻手指、手指黏在一起、出現兩隻右手。
生成文字背景任何表面出現似字母、數字或者筆劃的痕跡,模糊到讀不出都一樣作廢。
變化動作沒有在片內做完,畫面再靚,維持一個狀態就是作廢。
雜物背景空無一物,或者散焦到甚麼都看不見。

寧願早一步作廢,重跑一次都只是 368 credits。

第五步:所有文字,一律留到後期才加

兩條成品畫面上見到的字,全部是後期加的。我們試過叫 Seedance 2.5 寫字,中文英文都一樣,每個 take 都出錯,prompt 怎樣寫都救不回來,所以價錢、offer、電話、WhatsApp、logo、字幕和尾卡,一律在 CapCut 或者 Canva 加。

尾卡放在片尾那片留白上面,最後才落字幕。Reels 多數靜音播放,字幕不能省,而且同樣要自己打。

尾卡是一張靜態圖,在 Canva 開 1080x1920 做好,蓋住最後兩秒,或者片完之後再停多兩秒,如果叫模型生成尾卡,品牌名多數會串錯。

第六步:加廣東話配音與 offer

一條沒有 offer 的片是 b-roll,不是廣告。畫面再好看,裡面也沒有一句叫人做任何事,回報就是零。

配音 Seedance 2.5 講得出廣東話,我們生成過一段廣東話旁白,聽得明,讀音正確,節奏比真人硬一點。想控制語言,要用官方那條公式,把語言、地區口音和語氣風格三樣寫在說話者和台詞前面,我們試過用一個香港人設要求普通話,出來接近廣東話,人設壓過了語言指令,同一個要求改用官方公式重寫,出來就是標準普通話。聲音部分另有一套語法,() 放音樂,<> 放音效,{} 放對白,【】 放字幕。

用真人聲還是生成聲?生成的廣東話夠應付 b-roll、內部 demo 和大量測試,但要出街的廣告,我們會自己用手機錄,因為那點硬節奏偏偏就落在講 offer 那一句上面,而那一句最輸不起。錄一段 voice memo 不用錢,兩分鐘就做好。

Offer 怎樣寫 一個夠具體的 offer 要答到四件事,客人得到甚麼、哪些人適用、幾時到期、怎樣領取,所以寫「首杯奶茶減 $5,8 月內有效,到店出示這條片」,客人就知道下一步要做甚麼,而寫「歡迎試試我們的奶茶」,客人甚麼都不會做。

同時要合乎 Meta 的廣告政策,誇大或者不切實際的效果、使用前後對比,以及任何暗示你知道觀眾個人狀況的寫法,Meta 都會拒絕,所以文案只寫產品,不要寫觀眾的身材、健康、財務或者身份。offer 裡面每一個數字都要真的兌現得到,倒數計時永遠不會完、「首 20 位」永遠不會滿,這兩種寫法審核都不會收。逼真的 AI 生成內容,Meta 要求廣告主自己標示,政治與社會議題廣告更加是明文規定,所以要在 Ads Manager 開啟那個選項,不要自己判斷這條規定與你無關。最後,offer 寫甚麼,落地頁就要看得到甚麼,審核會連 destination 一併看。

兩條用 Seedance 2.5 生成的 AI 廣告片

兩條片都是照上面六個步驟做出來,分開跑三條 8 秒再順序剪起,加上在 Canva 做的尾卡,一條廣告合共 1,104 credits,剪片再用大約二十分鐘。

我們當時分開跑,是因為還未確認一鏡到尾守不守得住連戲。試完之後知道守得住,所以這兩條片今日重做,會直接生成 20 秒一鏡到尾,920 credits,省回那 184 和二十分鐘剪片。分開跑仍然有用,用途是想同一個場景出幾個版本落廣告。

產品參考圖用一張菠蘿油相,場地參考圖用一張茶記內部相。三段分別是一隻手把牛油放上菠蘿油、鏡頭推近影住枱上那舊牛油、最後拉開見到菠蘿油連一杯奶茶擺在枱面。連戲其實對不上:第一段是淺色雲石枱面配光管的冷光,第二段推得太近,背景只剩一片啡色的散焦,第三段變成深啡色的玻璃面枱,玻璃下面壓着一張餐牌,後面卡座還坐滿了客人。三段駁得起,靠的只是那個菠蘿油由頭到尾都一樣。還有兩處按我們自己那張檢查表是要作廢的:第三段玻璃枱下面壓着的餐牌,一行行是似字非字的筆劃;同一格後面幾張客人的臉也入了鏡。真要出街,這一段應該重跑,或者剪走那半格。

這條的產品參考圖是一張車相,場地參考圖是一張汽車美容工場相,三段是高壓水槍沖車頭、手拿洗車手套抹車身泡沫、最後一段抹乾淨的車身。連戲的做法跟上面一樣,但車徽避不到,車頭的廠徽連公牛都清楚出現。至於保險桿上那行字,是模型自己作出來的,逐格還會變字,所以由參考圖漏出來的是徽章,那行字則是生成文字失敗的又一次示範。漆色也走了,第一段是亮藍,之後兩段變成深藍,這一項在檢查表上同樣算作廢。背景那塊綠色告示牌上面的白字,一樣是模型自己作的。

這兩條片的水準去不到參賽級數,但一間香港中小企當日想到,當日就可以配一個細預算把它放出去。

你想試同一套做法,可以用自己那張產品相和舖頭相,在 Lumina 跑一條 8 秒,368 credits 就知道這條路走不走得通。

十條 AI 生成影片 prompt 讓你直接複製

下面十條是一個參考庫,每條都是獨立的 8 秒片,本身不接成一條廣告,你可以單獨用,也可以抽一條做三鏡廣告的中段,開場和收尾自己再寫。每一條的排法一樣,先是影片,然後一句做得到、一句穿崩位,最後是可以直接複製的完整 prompt。

十個行業各一條 AI 生成影片,全部用同一套方法做出來

十條用的設定都一樣:9:16・720p・8 秒・text-to-video・368 credits。 只有第 4 條例外,它用 image-to-video 配一張真實產品相。

1. 茶餐廳:沖絲襪奶茶

做得到: 十條裡面我們最滿意這一條,絲襪袋染滿茶漬,鋼壺燒到變色,防火膠板上留著杯印和水漬,旁邊疊著搪瓷杯和一塊濕抹布,看上去就是一間做了二十年的店舖。

穿崩位: 第一個 take 就收貨。

按此展開完整 prompt
Two hands pouring hot tea through a cloth tea sock held in a brass ring into a
dented stainless steel pot, the dark amber stream already falling in the first
frame, foam building on the surface as the pot fills and thick steam rolling up
across the frame. The counter is a worn cream laminate top with old cup rings
and a wet patch, a stack of chipped white enamel cups pushed to one side, a
damp stained cloth bunched next to them, a scratched steel tray holding two
more cups, and behind that plain grubby wall tiles with darkened grout, all of
it slightly out of focus. Flat cool overhead light from a bare tube fitting with
a warm patch of daylight from the left. The camera holds one fixed handheld
frame with small natural drift for the full duration, no push in and no other
camera movement. Photorealistic handheld phone footage, true to life texture,
mildly overexposed highlights, no cinematic grading. Vertical 9:16 composition,
no faces and no people beyond hands and forearms, hands have five fingers in
every frame, no text, letters, numbers, logos, signage, menus, price cards,
labels or printed packaging anywhere, every surface in frame blank and unmarked,
no machines or electronic equipment, shapes stay stable with no warping and no
flicker.

2. 咖啡店:拉花

做得到: 奶柱落到咖啡上,拉花的紋路一層層推開都乾淨,crema 的顏色對得上真機沖出來的樣子,鋼枱上的花痕和散落的咖啡粉也夠實在。

穿崩位: 跑一次就用得。

按此展開完整 prompt
Two hands pouring steamed milk from a small stainless jug into a plain white cup
of espresso, the pour already running in the first frame, the crema splitting
and a leaf pattern forming and settling as the cup fills to the rim. Shot
straight down onto a scratched stainless steel bar top scattered with loose
coffee grounds, a damp brown stained bar towel crumpled at the edge, a stack of
plain white saucers, a spoon lying in a small puddle, and water beads across the
metal. Warm side light from a window on the left, cool fill from above. The
camera stays in one fixed overhead frame with a small handheld sway for the full
duration, no push in and no other camera movement. Photorealistic handheld phone
footage, true to life texture, no cinematic grading. Vertical 9:16 composition,
no faces and no people beyond hands and forearms, hands have five fingers in
every frame, no text, letters, numbers, logos, signage, labels or printed
packaging anywhere, every surface in frame blank and unmarked, no machines or
electronic equipment in view, shapes stay stable with no warping and no flicker.

3. 花店:包花

做得到: 這條放大看都站得住,工作枱上散著花枝碎、麻繩碎和水漬,麻繩一收緊,牛皮紙就跟真紙一樣皺下去。

穿崩位: 一次過,沒有要報的。

按此展開完整 prompt
Two hands wrapping a bunch of white and pale pink flowers in brown kraft paper
on a battered wooden work bench, the hands already folding the paper in the
first frame, then drawing a length of natural twine around the stems and pulling
it tight so the paper crushes and creases sharply and two cut leaves fall onto
the bench. The bench is covered in cut stems, leaf trimmings, short offcuts of
twine and dark wet patches, with a grey plastic bucket of water half in frame
and a pair of plain scissors lying open beside it. Soft daylight from a window
on the right, cool and slightly grey. The camera makes one slow tilt down from
the flower heads to the hands over the full duration, no other camera movement.
Photorealistic handheld phone footage, true to life texture, no cinematic
grading. Vertical 9:16 composition, no faces and no people beyond hands and
forearms, hands have five fingers in every frame, no text, letters, numbers,
logos, signage, labels, printed paper or printed packaging anywhere, the kraft
paper is completely blank, every surface in frame unmarked, shapes stay stable
with no warping and no flicker.

4. 護膚品:配真實產品相

做得到: 這條可以直接落廣告用,label 上的 dermalogica 逐個字母正確,註冊商標符號在,連下面三行細字都與參考相對得上,要把真實 label 帶進生成影片,我們試過的方法之中只有這一個做到。

穿崩位: 沒有,不過這條跑了兩次。第一個 take 有一隻手遮住 label,樽身又有一半出了框,那張參考相等於白費,我們改寫 prompt,把樽身釘在畫面中央,並且寫明 label 不可以被遮,第二個 take 就乾淨。要帶真 label 入片的話,預兩個 take 會安全一點。

按此展開完整 prompt
@image 1 is the product: reproduce this exact bottle and its exact label
faithfully, do not redraw, restyle, re-space or re-colour it. Two hands lift the
bottle from a cluttered bathroom shelf and twist the dropper open in the first
frame, then squeeze one drop of clear serum onto the back of the other hand
where it lands and slowly spreads and catches the light. The shelf is crowded
with a damp folded towel, a plain unmarked ceramic cup holding a toothbrush, a
hair tie, a small dish with water pooled in it, and beyond it fogged tiles with
water spots and a mirror edge, all softly out of focus. Soft daylight from one
side, slightly cool. The camera holds one fixed handheld frame with small
natural drift for the full duration, no push in and no other camera movement.
Photorealistic handheld phone footage, true to life skin texture with visible
pores, no cinematic grading. Vertical 9:16 composition, no faces and no people
beyond hands and forearms, hands have five fingers in every frame, no text,
letters, numbers, logos or labels anywhere except the label carried in from the
reference image, every other surface in frame blank and unmarked, shapes stay
stable with no warping and no flicker.

上面幾條都可以直接複製,你想試自己那一行,可以在 Lumina 開個帳戶,把行業字眼和產品相換成自己的,其餘照跑。

5. 寵物美容:沖水(作廢)

做得到: 水的物理做得對,水流沖落之後濕毛貼服變深色。

穿崩位: 出事的是隻狗。頭轉開又切出了畫面,剩下一隻又長又尖的耳朵接住一條長頸,看上去似鹿多過似狗,泡沫的質感亦糊。這條作廢,但照樣放上來。我們沒有再跑第二次,所以只知道這一條救不回。

按此展開完整 prompt
Two hands lathering white soap suds along the wet back and shoulder of a small
brown dog standing in a shallow stainless steel tub, then lifting a plain
plastic jug and pouring warm water over the fur so the suds slide off in sheets
and the wet coat flattens and darkens. Only the dog's back, shoulder and one ear
are in frame and its head is turned away from the camera. The tub sits on a
scuffed tiled floor with puddles and stray wet fur, a soaked towel draped over
the tub rim, a plain rubber mat and an overturned plastic basin beside it.
Bright flat daylight from a window on the left. The camera holds one fixed low
handheld frame with small natural drift for the full duration, no push in and no
other camera movement. Photorealistic handheld phone footage, true to life wet
fur texture, no cinematic grading. Vertical 9:16 composition, no human faces and
no people beyond hands and forearms, hands have five fingers in every frame, no
clippers, dryers, machines or electronic equipment anywhere, no text, letters,
numbers, logos, signage, labels or printed packaging anywhere, every surface in
frame blank and unmarked, the dog keeps a stable anatomy with four legs and no
warping and no flicker.

6. 服裝網店:布料

做得到: 布料的織紋見到麻紗的粗結,下擺散開之後擺動幾下就停得穩,沒有卡在半空定格。

穿崩位: 摺完之後又打開。雙手把恤衫摺好按平,跟着整件提起攤開,八秒之內做完一件事又還原,結尾和開場一模一樣。這條正正犯了上面第三條規則,動作要有後果,而這條的後果是甚麼都沒有發生。次要的是背景那疊衣物仍然有淡淡的假商標痕跡,prompt 寫過禁止都照出,剪走那個角落或者選一格背景更散焦的畫面就可以。

按此展開完整 prompt
Two hands press and smooth a folded oatmeal linen shirt on a worn wooden table
in the first frame, then pick it up by the shoulders and let it fall open so the
fabric unfurls downward, the creases releasing and the hem swinging and settling.
The table is crowded with a leaning stack of folded garments in muted colours,
a loose thread, a crumpled sheet of plain tissue paper, two bare wooden hangers
and an open grey plastic crate at the edge, all slightly out of focus. Soft
daylight from a large window on the right, cool and even. The camera holds one
fixed handheld frame with small natural drift while the fabric moves through it,
no push in and no other camera movement. Photorealistic handheld phone footage,
true to life woven fabric texture with visible slubs, no cinematic grading.
Vertical 9:16 composition, no faces and no people beyond hands and forearms,
hands have five fingers in every frame, all fabric completely plain with no
print, no pattern, no tags and no stitching text, no text, letters, numbers,
logos, labels, tape measures or printed packaging anywhere, every surface in
frame blank and unmarked, shapes stay stable with no warping and no flicker.

7. 汽車美容:抹泡

做得到: 這條同時示範了第三條規則,泡沫成片滑走,抹乾淨那一塊立即反出天空,下面的水泥地積著水。

穿崩位: 沒有,因為構圖根本不給它機會,車徽、水箱罩、車頭燈和車輪都不入鏡。你籠統描述一部車,模型就會搬真實車廠的設計語言出來,所以要靠取景,把會露出品牌身份的位置避開。

按此展開完整 prompt
A hand in a soaked wash mitt drags across a wet deep blue painted car panel
already covered in white foam in the first frame, the foam sheeting away behind
the mitt and clear water running down while a sharp mirror reflection of the sky
opens up on the cleaned paint. Framed tight on a plain curved painted panel with
no badge, no grille, no lights, no window and no wheel in view. Below and behind,
soft focus wet concrete with puddles, a black bucket with a coiled hose beside
it and a damp folded towel over the bucket rim. Overcast daylight, cool and
diffuse, with one bright soft highlight sliding along the paint. The camera makes
one slow lateral follow alongside the panel for the full duration, no push in and
no other camera movement. Photorealistic handheld phone footage, true to life
water and paint texture, no cinematic grading. Vertical 9:16 composition, no
faces and no people beyond a hand and forearm, the hand has five fingers in every
frame, no text, letters, numbers, logos, badges, emblems, signage or printed
packaging anywhere, every surface in frame blank and unmarked, no machines or
electronic equipment, shapes stay stable with no warping and no flicker.

8. 網店包裝:封箱

做得到: 這條一次就跑到,膠紙拉出來貼上去,掌根一壓就平,紙箱到最後真的封好,牛皮紙和膠紙的質感都做得出來。

穿崩位: 沒有值得報的。

按此展開完整 prompt
Two hands fold the flaps of a plain kraft cardboard box shut in the first frame,
then pull a strip of clear packing tape across the seam and press it down firmly
with the heel of the palm so the tape flattens and the box is sealed. Shot
straight down onto a scarred wooden desk covered in working mess: torn offcuts
of bubble wrap, a crumpled ball of plain paper, three more flat kraft boxes
stacked unevenly, a pair of scissors and a scattering of paper dust. Warm
tungsten light from above with a cooler daylight edge from the left. The camera
stays in one fixed overhead frame with a small handheld sway for the full
duration, no push in and no other camera movement. Photorealistic handheld phone
footage, true to life cardboard and tape texture, no cinematic grading. Vertical
9:16 composition, no faces and no people beyond hands and forearms, hands have
five fingers in every frame, the boxes and tape are completely blank with no
print, no text, letters, numbers, logos, barcodes, labels, address slips or
printed packaging anywhere, every surface in frame unmarked, no machines or
electronic equipment, shapes stay stable with no warping and no flicker.

9. 珠寶配飾:倒鏈

做得到: 幼鏈成一條線連續落下,在絨布上捲成鬆散一堆,落到布面那刻反出細碎的光點,金屬跟後面那塊深色布分得清。

穿崩位: 沒有。prompt 禁止了寶石切面、印記和雕刻,而這幾樣都是模型會嘗試寫字的表面,所以這條乾淨。

按此展開完整 prompt
A fine yellow metal chain already pouring from between two fingers in the first
frame, falling in a thin continuous line onto dark crumpled velvet cloth where it
coils into a loose pile, catching small bright points of light as it lands. The
velvet lies on a worn wooden tray with a shallow ceramic dish of loose plain
rings pushed to one side, a soft grey polishing cloth bunched next to it and fine
dust visible on the wood. A single warm lamp from the upper left with deep shadow
on the right. The camera holds one fixed close handheld frame with small natural
drift for the full duration, no push in and no other camera movement.
Photorealistic handheld phone footage, true to life metal and velvet texture, no
cinematic grading. Vertical 9:16 composition, no faces and no people beyond hands
and forearms, hands have five fingers in every frame, no gemstones with cut
detail, no hallmarks, no engraving, no text, letters, numbers, logos, boxes,
labels or printed packaging anywhere, every surface in frame blank and unmarked,
shapes stay stable with no warping and no flicker.

10. 按摩推拿:落手

做得到: 油倒入掌心,雙掌搓熱之後推出去,手本身正常,五隻手指齊全,燭光、蒸氣和後面捲好的毛巾都在。

穿崩位: 枱上根本沒有人,布蓋著的是一張平枱,見不到肩也見不到背,雙手就在一塊平布上推油。要拍身體上的動作,主體必須寫進 prompt,模型不會自己補一個人出來。

按此展開完整 prompt
Two hands tip a small ceramic bowl and pour a little oil into one palm in the
first frame, rub the palms together, then press slowly and firmly along a
client's shoulder and upper back over a folded grey towel so the towel creases
and gathers under the pressure. The client lies face down with the head turned
fully away from the camera and out of focus. Around the table are rolled damp
towels stacked unevenly on a low wooden shelf, two burning candles, a folded
linen sheet slipping off the edge and a worn wooden floor below. Warm
candlelight from the right with soft shadow, fine steam drifting. The camera
makes one slow lateral drift along the client's back for the full duration, no
push in and no other camera movement. Photorealistic handheld phone footage, true
to life skin and fabric texture, no cinematic grading. Vertical 9:16 composition,
no visible faces, hands have five fingers in every frame, no medical or
professional equipment of any kind, props limited to towels, candles, bowls and
furniture, no text, letters, numbers, logos, signage, labels or printed packaging
anywhere, every surface in frame blank and unmarked, shapes stay stable with no
warping and no flicker.

AI 生成影片做不到的幾件事

之前那二十條測試片試出來的限制,這一輪同樣遇到。

模型寫不到字,一個四個字母的品牌名我們試了兩次都串錯,其中一次的 prompt 已經逐個字母拼給它,出來還是錯。中文招牌更差,交出來是似字非字的筆劃,同一個 take 裡面還會逐格變形,所以我們之後索性不讓帶字的表面入鏡,文字留到後期才加。

商標不用你提,模型會自己加上去。我們只是籠統形容一件產品,一個字都沒有提過品牌,但那個顏色組合剛好屬於一個著名品牌,模型就把整套設計語言搬了出來;上載一張真實品牌物件的參考相就更加直接,徽章會原原本本還原出來,我們在 prompt 裡面寫過幾句禁令,一句都攔不住。既然參考圖壓得過 prompt,真正控制到畫面的其實是取景,還有你上載了哪一張相。

專業儀器和醫療設備我們做不出來,要求過幾次,交出來的東西世上根本不存在,健身室本來在名單上,就是因為器材認不出來而剔走。

上載真人相會被平台拒絕,我們試過一張逼真的人臉,平台直接回一句「The input image may contain real people」,所以想做一個固定的 AI 代言人,這條路走不通。

我們那個 plan 在 2026 年 8 月跑的時候最高只出到 720p,所以上載之前會先放大,因為放在原生拍攝的 Reels 中間,觀眾會先嫌畫質差,然後才想到這是 AI。

上面這些限制,只有參考圖繞得過,前面兩條完成品廣告就是這樣做出來,其他方法我們試過都不行。每一項失敗的截圖和逐條對證,都放在Seedance 2.5 二十條測試片

AI 生成影片值不值得用

這一輪我們跑了十條單片,九條可以用,剩下一條因為動物結構作廢,我們沒有再重跑。兩條 20 秒完成品廣告每條用 1,104 credits,另外再花 1,380 credits 跑一條 30 秒單鏡測試去比較長片和短片,連重跑計在內,這一輪大約用了 8,200 credits。

一條片像不像一盤真實的香港生意,差別在枱面上有多少雜物、有沒有一雙手在做事、那件事有沒有做完、鏡頭有沒有亂動,形容風格的字眼依然幫不到甚麼;至於三條片像不像同一盤生意,我們試過的方法之中只有參考圖做得到。這兩樣做對,尾段放一個 offer 就有一條廣告。

想看整個模型的拆解和每一張失敗截圖,可以去Seedance 2.5 完整實測;想自己動手,先開個 Lumina 帳戶,準備兩張相,再寫三條 prompt 就可以開始。

常見問題(FAQ)

一條 20 秒的完成品 AI 廣告要多少錢? 一鏡到尾生成是 920 credits,拆成三條 8 秒剪起就是 1,104,每條 368。尾卡和畫面上的文字在後期加,不用額外生成費。

一條 8 秒 AI 廣告片多少錢? 368 credits,720p,收費按秒線性計,每秒 46 credits,畫面比例不影響價錢。

兩張參考圖是不是一定要? 想條片認得出是你那間舖就要,分開跑三條要,一鏡到尾同樣要,因為產品和場地都靠它固定。純文字 prompt 每次生成的場景都不一樣。只出一條單獨的片,不掛也可以。

一鏡到尾,是不是好過三條 8 秒? 連戲它守得住,由頭到尾同一部車、同一種光,轉場由它自己接。價錢也平一點,一鏡到尾 20 秒收 920 credits,三條 8 秒剪成是 1,104。分開跑的主要好處是改動便宜,重跑一條 368,重跑成條 20 秒 920。想同一個場景出幾個版本落廣告就分開跑,prompt 寫得穩就一鏡到尾。

可不可以用有品牌的產品相做參考圖? 品牌是你自己的就可以。模型會照著你上載的相還原,prompt 寫明不要 logo 也攔不住,我們拿一張真車相試過,車廠徽章清楚出現在畫面上。

可不可以把自己產品的真 label 放進 AI 片? 可以,用 image-to-video 掛一張產品相上去,我們那條護膚品片的 dermalogica label 逐個字母正確,註冊符號和細字都在。交給模型自己寫字則做不到。

Seedance 2.5 講不講到廣東話? 講得到,聽起來也明白,只是節奏比真人硬一點,貼近讀稿多過對話。要控制語言,把語言、口音和語氣寫在說話者和台詞前面,例如 Dialogue language: Cantonese (香港廣東話). The owner says in a warm natural Hong Kong accent: {台詞}。不寫清楚它會自己配一個人設。要出街的廣告我們用手機錄真人聲,生成的聲音留給 b-roll 和 demo。

燈光明明打得夠好,為甚麼出來仍然假? 多數是場景太乾淨。為了避開文字失敗,我們一度把背景清空,結果枱面上一件東西都沒有,燈光打得再好都假。把雜物逐件寫進 prompt,並且選不帶字的。

哪些香港行業最適合? 分界線是鏡頭裡要不要一個人的身體。靠手、水、布、紙、食物和蒸氣去表達的都做得到,我們試過茶餐廳、花店、汽車美容、網店包裝和珠寶,第一個 take 就用得。要影住真人接受服務的行業,我們暫時做不到,那條按摩片枱上根本沒有人,布蓋著的是一張平枱。脊醫和美容療程同類的我們未試過。靠器材和招牌的行業,我們試過健身室,做不出來。

寵物美容那條為甚麼作廢? 狗的頭轉開又切出了畫面,剩下一隻長耳接住一條長頸,看上去似鹿多過似狗。我們沒有再跑第二次,所以只知道這一條救不回。

Prompt 用中文還是英文寫? 都可以。同一條 brief 我們中英各跑一次,質素分不出高低。

Seedance 2.5 寫不寫到中文招牌? 寫不到,中英文一樣,做法是讓招牌留白,字在剪片時才疊上去。

應該預多少個 take? 一般預一個,帶真 label 或者有活生生動物的預兩個,如果是三鏡廣告,中段那個鏡頭我們會多預一次。

Frankie Chan

關於作者

Frankie Chan · Kick Ads 聯合創辦人

Frankie 是 ex-Googler 及廣告策略師,自 2017 年起為香港及馬來西亞的電商及 Lead Generation 企業管理 Google Ads 及 Meta Ads,專注策略、報告及增長規劃。

想由專家檢視你的帳戶?

做一次 Kick Ads 的 Paid Media Health Check:沒有 sales pitch,只有實話。

免費廣告 Health Check