SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

📘 SlideCast Studio Official Manual

💡 This manual is also available on the following website. It is formatted for better readability, so please be sure to check it out!

Update Information
▪️ August 8, 2026: Content significantly updated to match the latest version, v3.4

SlideCast Studio Welcome to the official manual for.

SlideCast Studio is a slide video creation tool that I 'magically' modified after seeing Majin-san's note. I was only able to create this tool because Majin-san made the prompt public.

And here is SlideCast Studio v1.0, which I created by 'magically' modifying that prompt and turning it into a GAS application.

And this SlideCast Studio v2 and beyond is a major update to that SlideCast Studio v1.0.
Click here to use it👇


Chapter 1: Introduction & Key Features

🎬 What is SlideCast Studio?

It is an application that allows you to load PDFs, images, and videos into your browser, add AI narration audio, and turn them directly into videos. No installation is required; you can complete your video entirely within your browser.

✨ Main Features

  • 🤖 AI Script Generation: Gemini integration allows you to automatically create scripts just by providing your materials

  • 🔊 AI Voice Synthesis: Narration selectable from 30 different voices (Gemini TTS)

  • 🎭 Character Performance: Lip-sync, facial expressions, poses/gestures, and breathing motions

  • 🎞️ Slide Animation (Ken Burns): Add movement to still images with pan & zoom

  • 🎬 Video Slides: Insert MP4/WebM as slides (playback range trimming available)

  • 🧱 Blank Slides: Create titles, chapter breaks, and conversation scenes without source materials

  • 🎨 Design Settings: Gradients, borders, shadows, rounded corners, badges, logos, and captions

  • 🎵 BGM (Up to 10 tracks): Switch BGM per slide + crossfade

  • 🎤 Microphone Recording: Record within the app (excluding GAS version)

  • 🔔 Sound Effects (SFX) / Text Scrolling Sounds: Transition sounds, retro-style text sounds for silent dialogue

  • 🕐 In-slide timeline: Adjust audio, video, and intervals by the second

  • 📚 Book-style PDF: A4 portrait format with integrated slide images and subtitles

  • 🌐 UI language: Supports Japanese and English

  • 💾 Export formats: MP4, WebM, PPTX, PDF, PNG, JPEG


Chapter 2: System Requirements and Recommended Specs

🌐 Recommended Browsers

  • ✅ Google Chrome (PC version, latest) ← Highly recommended

  • ✅ Microsoft Edge (latest version)

  • ⚠️ Smartphones and tablets are not recommended

💻 Guidelines for number of slides

Video rendering uses your computer's browser memory.

For 8GB memory

  • 720p: Approx. 20–40 slides

  • 1080p: Approx. 10–20 slides

  • 1440p: Not recommended

For 16GB memory (recommended)

  • 720p: Approx. 50–100 slides

  • 1080p: Approx. 25–50 slides

  • 1440p: Approx. 10–20 slides

For 32GB memory

  • 1080p: Approx. 50–100 slides

  • 1440p: Approx. 20–40 slides

There are three output resolution options: 720p / 1080p / 1440p (for 16:9, these are 1280×720 / 1920×1080 / 2560×1440).

Other Recommended Specifications

  • OS: Windows 11 or later, or macOS 13 or later

  • CPU: 6 cores or more recommended

  • Storage: SSD recommended

🐢 If performance feels sluggish

  1. Lower the output resolution from 1440p to 1080p or lower

  2. Lower the video quality from "High" to "Standard" or "Low"

  3. Lower the import quality of your materials and re-import them (Chapter 7)

  4. Clean up unnecessary projects or assets

  5. Restart your browser

  6. Split long projects into smaller parts

🫨 If the exported video feels choppy

Video exporting uses browser memory, so please pause other tasks. Closing unnecessary tabs will also improve the quality of the exported video.

In v3.4, recommendations for stable exporting will be displayed before exporting (you can hide this next time. To show it again, go to "Settings → Basic → Stabilize Video Export").


Chapter 3: Let's Start with Bundle Builder (Highly Recommended)

🏗️ What is SlideCast Studio Bundle Builder?

Bundle Builder is a dedicated auxiliary tool that allows you to start video production without setting up an API key (a standalone app that runs on Google Gemini Canvas). It guides you through the process using a step-by-step wizard, from script generation to AI voice synthesis and ZIP output.

You can open it from the Bundle Builder card (with the "Start Here" badge) on the dashboard.

🌟 Key Benefits of Bundle Builder

  • No API key setup required: You can start with standard models without needing to obtain your own Gemini API key.

  • Wizard format prevents confusion: Completed in 5 steps.

  • Supports Japanese and English: Switch using the 🌐 button at the top right of the header.

📋 5 Steps of Bundle Builder

Step 1: Input
Choose how to prepare your materials.

  • Generate materials on the spot (v3.3 or later): Specify the theme, target audience, number of slides, art style, etc., and have the AI create the slide images themselves. You can regenerate the entire set or individual slides after generation.

  • Upload materials: Import PNG / JPEG / PDF / ZIP / folders.

You can also set the title to be used for the output filename here. If you plan to use microphone recording, it is smoother to complete microphone permission on this screen.

Step 2: Script Generation
Generate after deciding on the target, script style, length, number of speakers (up to 6), subtitle line breaks, subtitle layout, and reading error prevention (furigana addition).

After generation, you can freely edit the speakers, tone, body text, and subtitles in the Script Editor. You can also add/delete scripts, copy/import JSON, and revert to the previous script.

Step 3: Audio Generation
Automatically extract speakers from the script, set voice models for each speaker, and generate in bulk. Supports adjusting the number of simultaneous generations, stopping, and regenerating only incomplete parts. You can also record with a microphone and replace it (with trimming before and after).

Step 4: Verification
Verifies 4 items: script check (empty text, sequential numbering), speaker consistency, audio links, and slide references.

Step 5: ZIP Output
Download a ZIP containing the slide materials, script, and audio.

📥 Importing into SlideCast Studio

Simply drop or select the output ZIP in Create from materials on the dashboard. Then, refine the design and presentation in the editor and export.


Chapter 4 Data Storage and Security

💾 Where is it saved?

All project data is saved "within your browser (IndexedDB)." No uploading to external servers occurs.

⚠️ Cases where data is lost

  • When you clear your browser's "browsing history/cookies"

  • When you uninstall the browser

  • When you close Incognito (private) mode

🆘 Once data is lost, it cannot be restored by the app. Regular backups are the only solution.

🔐 Precautions when using an API key

  • Avoid entering it on shared PCs or at internet cafes at all costs

  • The API key is not saved in project data; it is stored only within the browser

  • If you turn on "Do not save to this device (session only)", it will be deleted as soon as you close the tab/browser

  • We strongly recommend setting usage limits and alerts in the Google Cloud Console

  • If your API key is leaked and misused, there is a risk of high charges


Chapter 5 Project Management (Dashboard)

📊 Checking data usage

The estimated data usage within the browser is displayed at the top of the dashboard. If you are approaching the limit, delete unnecessary projects or back them up and move them elsewhere.

➕ How to create a project

Create from materials (recommended)
Select or drag and drop PDF, PNG, JPEG, or ZIP files (or folders). Multi-page materials may take a few minutes.

In v3.4, you can choose the import image quality before importing (Chapter 7).

Create an empty project
Use this if you want to start from scratch. You can add materials later.

🗂️ Project operations

  • 📋 Duplicate: Click the copy icon

  • ✏️ Rename: Click the title or the edit icon

  • ☁️ Backup Save: Save icon → Download file

  • 🗑️ Delete: Trash icon (cannot be restored after deletion)

The card displays actual data estimate and estimated capacity for individual backups.

💾 Backup and Restore

From the "Save Data" button, you can select items to include and export as a ZIP.

The selectable items are as follows:

  • Environment Settings

  • Characters

  • Custom Presets

  • Color Palettes

  • Media Library

  • Projects (can be selected individually for saving)

The estimated backup size is displayed, and it warns in three levels: Safe/Caution/Danger. At the danger level, restoration may fail, so individual backups are recommended for projects.

Restore (Import)
"Data Restore" button at the top of the dashboard → Select or drop a JSON or ZIP file. You can choose the action to take if a character or project with the same name is found.

  • Skip duplicates

  • Overwrite duplicates

  • Add with a different name

🗑️ Category-based Data Initialization

In "Data Initialization," you can select the categories to delete by checking them.

  • Project

  • Asset Library (Common assets such as images, videos, and BGM)

  • App Settings (Characters, Presets, Color Palettes, Gem URL)

  • API Keys and other local settings

A button to open a backup is provided before deletion.Deleted data cannot be recovered.


Chapter 6: Basic Editor Screen Operations

🖥️ The 4 Screen Areas

① Header Area (Top)
Return to Dashboard button, project name (click to change), ⚙️ Settings button, statistics (estimated playback time, estimated file size), script/bulk generation/slide output/video export buttons, global playback button, save status (Unsaved/Saved)

② Sidebar (Left)
List of slide thumbnails. Drag and drop to reorder, drop images/videos/PDFs/ZIPs to add in bulk, use the hover menu to change images, generate TTS, duplicate, or delete. **Transition effect boundary badges (Global/Individual)** are displayed between thumbnails. The width can be adjusted by dragging.

③ Preview Area (Center)
Displays the canvas in real-time. Select the preview background from Light/Medium/Dark. Drag the border to adjust the area size. Includes playback controls (▶ / ⏸ / seek bar) and a "Timing Adjustment" button.

④ Script Editor (Bottom)
Enter dialogue and subtitles for each slide. This is also where you manage interactions between multiple characters, acting (expressions/poses) settings, and the slide timeline.

🟢 Easy Mode (v3.0 and later)

For those who feel there are too many settings, we have provided an"Easy" switch. When turned ON, only the minimum necessary settings are displayed, and others are automatically configured with recommended content.

The following items can be adjusted in Easy Mode:

  • Screen Size (Landscape/Portrait/Square)

  • Subtitles (Display ON/OFF, Top-Bottom Split/Overlay, 4 design types, font size)

  • Image Display (Fit to Screen/Fill)

  • Slide Size (100% / 95% / 90%)

  • Screen Background Color

  • Character (Name, Icon, Shape, Size, Speech Bubble, Display Side)

  • Opening/Ending Videos

  • BGM (Track/Volume)

If you want to make fine adjustments, turn the switch OFF to return to normal mode.

▶️ Playback Controls

  • Play current slide only: ▶ button at bottom center / click canvas / Space key

  • Play entire presentation: "Play All" in header → Full-screen playback including transitions, OP/ED, and BGM

  • Timing adjustment: Mode for enlarging the preview to fine-tune pauses and audio placement

⌨️ Shortcut Keys

  • Save: Ctrl / ⌘ + S

  • Play / Pause: Space

  • Previous / Next slide: ← / →

  • Delete slide: Delete

  • Duplicate slide: Ctrl / ⌘ + D

  • Undo: Ctrl / ⌘ + Z

  • Redo: Ctrl + Y / ⌘ + ⇧ + Z


Chapter 7: Adding and Managing Slides

➕ How to add slides

You can add slides from "Add Slide" at the bottom of the sidebar, or by dragging and dropping images, videos, PDFs, or ZIP files into the sidebar. All pages of a PDF will be automatically converted to images.

🖼️ Selecting Import Quality (v3.x)

You can select quality presets when importing materials.

  • Original Quality Priority: Import at high quality, capped at the original image resolution (PDFs are high-definition).

  • 4K: For 4K video assets. Uses more time and memory.

  • 2K: A balance of quality and speed.

  • Full HD: For standard YouTube use.

  • Standard: For general use. Slightly lighter.

  • High Speed: For review purposes. Fine text may appear soft.

⚠️ Resolution will not exceed that of the original image or PDF. Even if you select a high quality setting, the original data resolution is the upper limit.

🧱 Empty Slide (v3.4 New Feature)

Select "Add Slide" → "Add Empty Slide" to add a slide without any source material.

You can place subtitles, character icons, speech bubbles, captions, and badges on top of a project background (solid color, gradient, image, or video).

This is suitable for the following uses:

  • Title screens

  • Chapter breaks

  • Supplementary explanations

  • End cards

  • Dialogue-only scenes

If you set an image later, it will switch to a standard image slide. Ken Burns is an effect for images, so it is not displayed on empty slides. It supports previews, transitions, and slide image output.

🔀 Reordering and Hover Menu

Thumbnails can be reordered by dragging and dropping. You can perform the following operations from the menu displayed on hover:

  • 🖼️ Change Image: Change only the image while keeping the script as is.

  • 🔊 TTS Generation: Generate audio only for that slide

  • 📋 Duplicate: Copy the slide with the same image and script

  • 🗑️ Delete: Can be restored with Ctrl + Z

⏱️ Display time for static slides (Automatic/Manual)

You can switch the display time in the "Static Slide Settings" of the script editor.

  • Automatic: Automatically updates based on audio length or character count, and the "pause" before/after or between conversations

  • Manual: Specify the number of seconds directly

You can also set the motion (Ken Burns) and intensity for each slide individually here.

🎬 Inserting video slides

You can insert MP4/WebM files as slides (200MB / 10-minute limit).

There are three audio modes.

  • 🔇 Mute: Cut the video audio. BGM continues

  • 🎬 Video audio only: No narration. BGM is muted

  • 🎚️ Mix with BGM: Play both

In addition, you can configure the following settings:

  • Video volume

  • Playback range (trimming start and end positions)

  • Ducking: Automatically lower BGM while video is playing

  • Full Screen: Expand the video to the entire canvas (hides badges and captions).

  • Show Subtitles: Overlay subtitles on the video only when in full-screen mode.

  • Use "Preview with these settings" to check the audio output in advance.

💡 Ken Burns effect is not applied to video slides.


Chapter 8: Creating Scripts and Audio

🗂️ "Script" Tab (v3.x)

You can open all script-related operations from the "Script" button in the header.

  • Generate: Select conditions to generate a script on the spot and set it to the slides.

  • Load: Paste and import a script JSON.

  • Bundle: Load data from Bundle Builder.

  • Paste: Convert from text.

Generate tab allows you to specify the target, style, length, number of speakers, subtitle lines per page, subtitle layout, script language, furigana assistance, generation model, TTS method, and whether to perform research.

Instead of generating on the spot, you can also display the "Prompt for External AI", pass it along with your materials to ChatGPT / Claude / Gemini, etc., and import the returned JSON using the "Load" tab.

⚠️ Video slides are excluded from script generation (scripts are set in order only for static slides).

📝 Structure of Script Blocks (Dialogue)

You can add multiple "Script Blocks" to each slide.

TTS Manuscript (Left side)
The text that the AI will read aloud. For Gemini TTS, difficult kanji will be read accurately if you add furigana in hiragana or parentheses. Pauses are automatically inserted at commas (、).

Displayed Subtitles (Right side)
The captions displayed on the screen. If left blank, the TTS manuscript will be displayed as is. Use line breaks for simultaneous display, and empty lines (two line breaks) for page turns.

⏳ Individual Specification of Gaps

You can specify the following "gap" in seconds for each script block.

  • Lead-in Pause: Before the first script (default is Pre-Roll)

  • Pause: Waiting time until the next script (default is conversation interval)

  • Lead-out Pause: After the last script (default is Post-Roll)

Individual settings can be reverted to common settings with one click.

🔊 AI Voice Generation

Generate per block
Click the "Generate" button → A green "AI Generated" badge will appear in a few seconds.

Generate all slides at once
Select one of the following from "Batch Generate" in the header.

  • Generate all (overwrites existing audio)

  • Only ungenerated (retains existing audio)

Progress and remaining time are displayed, and you can stop at any time with "Emergency Stop".It will automatically stop if errors occur consecutively and generated audio will be retained.

⚠️ The Gemini free tier has limits not only on requests "per minute" but also "per day". If you reach the limit, you can also create audio in the Google AI Studio Playground and upload it.

🎭 Tone Specification (Control of emotion and speaking style)

When using Gemini TTS, you can specify the speaking style in the "Tone" field of each script block. Enter things like "Calm explanatory tone, slightly slow," "Bright and energetic," or "As if surprised."

🎤 Using External Audio

You can drag and drop MP3/WAV files into the script block or select files using the "Load" button.

🎙️ Microphone Recording

You can record, listen, and adopt audio (with trimming) using the "Record" button on each line of the script.

ℹ️ The record button is not displayed in the GAS version. Also, since the local HTML version is not intended to use the TTS function, only BGM settings are displayed for audio settings (recording and audio file loading can be used).


Chapter 9: Timeline and Timing Adjustment within Slides (v3.x)

In the script editor's "Timeline", you can check and adjust the contents of that slide on a time axis.

There are three tracks.

  • Video: Still images/videos (static displays before video playback are also shown)

  • Audio: Audio for each script block

  • Interval: Waiting time before/after and between conversations

Adjustments that can be made by dragging are as follows:

  • Audio start position

  • Video start timing

  • Still image display duration

  • Slide end point (moving this switches the still image to manual duration)

Display duration can be identified by the "Auto Duration" or "Manual Duration" badges, and you can revert at any time using "Revert to Auto." If audio clips overlap, an "Overlap" badge will appear.

Pressing "Timing Adjustment" at the top of the preview will enlarge the preview for fine-tuning.


Chapter 10 Design & Production Settings

The settings modal is divided into the following tabs: Basic / Subtitles / Production / BGM / Badges / Telop / Audio / Character / OP/ED. Settings are saved automatically and applied immediately (only character editing requires clicking "Save" to apply).

🎨 Color Palette

By registering frequently used brand or character colors, you can maintain consistent color schemes across projects. You can register both solid colors and gradients.

⚙️ Basic Tab

  • Presets: Save your preferred settings with a name and recall them with one click (items that will change upon application are displayed in advance)

  • Video Size: 16:9 / 9:16 / 1:1

  • Video Preview Quality: Lightweight / Standard

  • Slide Image Placement: Show Entire (Reset) / Fill (Zoom), Zoom Ratio, Position, and Slide Rounded Corners

  • Background/Wallpaper: Solid color / Gradient / Background image / Background video (Background video loops without audio)

  • Video Export Stabilization: Whether to display recommendations before exporting

Background images and videos can be saved to the Library and recalled from other projects with a single click.

💬 Subtitle Tab

  • Subtitle Layout: Overlay / Split top-bottom / No subtitles

  • Auto Display Duration: Display duration for subtitles without audio (automatically calculated based on seconds per character)

  • Font and Size: Select from around 20 fonts using a modal that allows you to compare their actual appearance. You can also specify font weight (thin/regular/bold) and italics.

  • Area Height: Height of the subtitle area (a warning will appear if any characters are based on this area height)

  • Display Position: Top-aligned / Center / Bottom-aligned, vertical margins, and horizontal margins

  • Display Animation: None / Fade in / Blur fade / Slide up / Slide from left / Pop / Typewriter

  • Style: Text color (solid/gradient), outline, shadow, and subtitle bar with opacity

In v3.4, while the subtitle settings are open, a "Subtitle Settings Preview" will appear on the canvas. You can check changes to font, size, color, outline, shadow, and area height in real-time.

🎞️ Effects Tab: Slide Animation (Ken Burns)

This effect adds slow panning and zooming to still images.

Available presets include: None, Zoom In (Center), Zoom Out (Center), Pan Left to Right, Pan Right to Left, Pan Top to Bottom, Zoom In Top-Left, Zoom In Bottom-Right, Picture Book Sway, Impact Zoom, Impact Shake, Pop Start, and Wave Sway.

  • Auto Variation: When ON, it automatically cycles through different effects for each slide.

  • Effect Intensity: Subtle / Standard / Strong

  • Hide background: Automatically zoom in slightly only if the edges are visible during movement

  • Individual settings per slide are also possible

⚠️ The editor preview is static. Please check the movement by playing or exporting. This does not apply to video slides.

🎞️ Effects Tab: Screen Transition Effects

  • Cut (no effect)

  • Cross Dissolve

  • Fade

  • Slide (Left/Right/Up)

  • Zoom In

  • Zoom Out/In

  • Circle Wipe

  • White Flash

  • Page Swipe

Transition time can be set in seconds. If you turn on "Apply to OP/ED video boundaries as well", the OP to first slide and last slide to ED transitions will also use the same effect.

🔀 Individual Settings per Boundary (v3.x)

By clicking the boundary badge ("Global"/"Individual") between slides in the sidebar, you can override settings only for that boundary.

  • Effect Type

  • Transition Time (Cut can also be specified)

  • Sound Effect (Follow global settings / OFF / Play global sound only here / Individual settings)

  • Individual Sound Effect Volume

You can return to the global settings with "Reset All".

🔔 Transition Sound Effects (SFX)

Built-in sounds include: Picture Book Page Turn / Thin Page Turn / Page Turn (Legacy) / Cinematic / Shimmer / Soft Whoosh / Paper Swipe.

You can also add custom sounds by uploading MP3s or similar files ({5MB limit}). Uploaded sounds are saved to your library and can be used in other projects.

🎮 Text-to-Speech Typing Sounds (v3.x)

This feature adds short, retro RPG-style sounds to dialogue lines that do not have audio.

  • Global ON/OFF, default narrator pitch (Extra Low / Low / Standard / High), and volume

  • ON/OFF, pitch, and volume can be overridden for each character

These sounds do not play during lines with actual audio, empty lines, or gaps between dialogue.

⏱️ Timing Adjustments (seconds)

  • Pre-Roll: Wait time from slide display until speech begins

  • Gap: Common wait time used between script lines

  • Post-Roll: Wait time from the end of speech until moving to the next slide

All of these can be overridden individually on the script block side (Chapter 8).

🏷️ Badge Tab

You can independently place logos or text badges in each of the four corners.

  • Create in App: Multi-line text, text alignment, text color/shadow/border, background (solid/gradient/transparent), shape (square/rounded), shadow, border

  • Upload Image: Place logos such as transparent PNGs. Adjust size and shadow

Badge images can also be saved to the library.

⚠️ Using this to hide watermarks from other services may violate terms of service.

📰 Telop Tab

This is a band-shaped decoration that flows across the screen. You can independently configure Top Telop / Bottom Telop.

  • Display Method: Fixed / Flowing / Slide Left

  • Assets: You can add multiple text and image items and have them flow together

  • Adjustment: Flow speed (px/sec), pause time, transition time, asset interval, grouping interval

  • Band: Style, color, opacity, text border/shadow, font


Chapter 11 Character Management

Character settings have changed significantly in v3.4. Items are organized into three groups: Character Assets/Movement / Layout / Design, and you can reset each group to its default values.

👤 Character Creation

Create from Settings → Character → "New Creation". Set the name, style (speech bubble/standard subtitle/band), font, text weight, italics, etc.

🖼️ Character Assets (3 Methods)

Images that can be registered are divided into three methods. Switch using the tabs at the top.

■ Registration Order Animation (Registration Order)
If you register multiple still images, they will switch in the order of registration while speaking. You can use this with as few as two images. You can reorder them by dragging the cards or using the arrow buttons.

■ 1-File Animation (Animated Image)
If you register one GIF / animated WebP / APNG file, that animation will play while speaking. You can also adjust the playback speed.

■ 6-Image Lip Sync (Audio Linked)
Switches between 6 still images according to the volume of the narration. The roles of the 6 images are fixed.

  1. Eyes Open, Mouth Closed (Idle)

  2. Eyes Open, Mouth Half-Open

  3. Eyes Open, Mouth Open

  4. Eyes Closed, Mouth Closed

  5. Eyes Closed, Mouth Half-Open

  6. Eyes Closed/Mouth Open

You can also replace a single image by clicking "Replace," which appears when you hover your mouse over an image.

Lip-sync interval is set to "Auto (Recommended)" by default. It estimates the speaking speed based on the audio duration and subtitle length for each line. Only adjust it using the "Manual" slider if you feel it doesn't match.

💡 Materials for the 3 modes are saved separately. Since previous materials are not deleted when you switch modes, you can easily try them out and switch back.

📥 How to add materials

  • Select multiple files

  • Select by folder

  • Drag & Drop (both files and folders are supported)

  • ZIP file (can be used in "Registration Order" and "Audio Sync")

The contents of the ZIP are loaded in alphabetical order by filename. It is recommended to number them, such as 01.png to 06.png, to ensure the correct order.

The limits are 50MB per ZIP / 200 images / 100MB total after extraction. The limit for a single-file animation is 12MB / 180 frames, and if it is too large, it will be automatically downsized as much as possible before loading (the size after downsizing will be notified).

🎭 MouthLoop Integration and Character Pack ZIP

MouthLoop is a separate app (running on Gemini Canvas) that automatically generates lip-sync materials from a single image. You can open it from the "Create materials with MouthLoop" link in the character settings.

  • 6 PNGs (ZIP) → Can be imported directly into "Audio Sync"

  • Animated WebP → Can be imported directly into "Animated Image"

  • Character Pack ZIP → In addition to lip-sync materials, you can register expressions and poses/gestures all at once

Character Pack ZIPs are automatically detected, and before loading, the number of images, extracted size, estimated storage capacity, included expressions, and included poses/gestures are displayed. A warning will appear if there is insufficient free space. Since regular image ZIPs and character packs are automatically distinguished, there is no need to switch the loading method.

😊 Expressions & Poses/Gestures (v3.4 New Feature)

Select the text in the displayed subtitles where you want to add acting, and the "Add Acting" button will appear. From there, you can set expressions and poses/gestures (multiple can be specified at once).

  • Make the character smile only during the "Thank you" part

  • Change facial expressions with surprising lines

  • Insert a pointing gesture in the middle of an explanation

Configured acting is listed below the script field and can be edited or removed later. If the target text is significantly rewritten, the acting settings may be automatically removed to prevent misalignment.

🪄 Auto-Acting (v3.4 New Feature)

When "Auto-Acting" is ON and you press "Auto-place acting" expressions and gestures will be placed according to the breaks in the lines.

  • Expressions/gestures to use (selectable per character)

  • Frequency: Low / Standard / High

  • Scope: This slide / All

Expressions matching built-in words like "Thank you" or "Too bad" are prioritized.Since it is not analyzed by generative AI, no additional API communication or API fees will be incurred.

Automatically placed acting is marked with an "Auto" tag, and you can delete only the auto-placed items with "Delete auto-placed items" (manual settings will remain). We recommend placing them automatically first and then manually correcting only the parts you are concerned about.

ℹ️ Acting timing is calculated from the position of the selected text and the length of the audio. It does not obtain word pronunciation timestamps via speech recognition, so it will not be perfectly synchronized word-by-word. Please use it for dramatic effect.

🫁 Basic Motion (Breathing/Speech Bounce)

You can also add movement to still character icons.

  • Slow breathing while silent

  • Speech bounce matched to the volume of the voice while speaking

Settings are in 4 levels: None / Weak / Standard / Strong. No new image assets are required, and the movement will be the same in preview and export.Existing characters remain at "None" so the appearance of your previous projects will not change.

ℹ️ The movement is only the vertical bobbing and slight stretching of the entire icon. It does not tilt or perform movements where the drawing itself changes, such as waving a hand.

📐 Layout (Size, Position, Alignment)

Icon size can be chosen from 2 methods.

  • Subtitle area based (Scale): Same behavior as before. Changing the "Area height" of the subtitles also changes the icon size

  • Screen-based (px): Does not link to subtitle settings. Even if the export resolution changes, the ratio relative to the screen is maintained.

Speech bubble height also has two methods.

  • Fit to subtitle area

  • Fit to text (expands/contracts based on number of lines)

In addition, the following adjustments can be made:

  • Vertical alignment with speech bubble (align center / align bottom)

  • Icon horizontal/vertical movement

  • Speech bubble horizontal position

  • Character-specific text size (global settings / per-character, based on 1080p)

💡 When you move the character's horizontal position, the speech bubble moves along with it by the same amount (v3.3 and later). Since the positional relationship is maintained, once you have arranged them neatly, you can move the whole set together.

🗨️ Design (Shape, Color, Border, Shadow)

  • Icon type: Circle / Original Art

  • Icon shadow

  • Speech bubble shape: Basic / Thought / Simple / Drop

  • Speech bubble color, shadow, border, and opacity

  • Band style color (when band style is selected)

📍 Standing Position

You can place them on the left or right side of the screen. For dialogue-based content, it looks natural to separate them left and right.

🔊 Voice Parameters

Gemini TTS (when using Gemini API key) — You can choose from 30 different voices.

Female / High Pitch: Achernar (Soft) / Laomedeia (Energetic) / Leda (Youthful) / Zephyr (Bright)

Female / Mid Pitch: Aoede (Light) / Autonoe (Bright) / Callirrhoe (Natural) / Despina (Smooth) / Erinome (Clear) / Gacrux (Mature) / Kore (Firm) / Pulcherrima (Positive) / Sulafat (Warm) / Vindemiatrix (Calm)

Male / Mid Pitch: Puck (Energetic) / Rasalgethi (Explanatory) / Sadaltager (Intellectual)

Male / Low-Mid Pitch: Achird (Friendly) / Alnilam (Firm) / Fenrir (High Energy) / Iapetus (Clear) / Orus (Firm) / Schedar (Uniform/Stable) / Umbriel (Natural) / Zubenelgenubi (Casual)

Male / Low Pitch: Algenib (Grit) / Algieba (Smooth) / Charon (Explanatory) / Enceladus (Breathy) / Sadachbia (Lively)

You can play a sample using the preview button (*Previewing also consumes API quota).

Google Cloud TTS (when using Google Cloud API key) — You can select Neural2 / WaveNet voices and adjust speed and pitch.

🎚️ Adjusting Set Voices (v3.x)

You can apply speed and pitch to generated, recorded, or uploaded audio after the fact.

  • After previewing, use "Apply to all lines" for batch re-conversion

  • "Revert to original audio" to restore the original sound

ℹ️ Previewing alone does not change the saved audio or the timeline.

🔄 Batch Character Replacement

Use this when you want to change Person A's lines to Person B. Select the source and target, then click "Execute" to replace all slides at once. You can also choose whether to keep the set audio.


Chapter 12: BGM Settings (Up to 10 tracks)

🎵 How to set BGM

Drop an MP3 or load it via "Add BGM". Default volume can be set, and you can override individual volume per track. If the video is longer than the BGM duration, it will automatically loop.

🎚️ Multiple BGM Tracks & Crossfade

You can register up to 10 tracks for BGM, and specify the "Playback start slide" for each. At the switching timing, crossfade is applied, and the fade duration can also be set in seconds.

Example: Intro uses calm BGM → Automatically switches to bright BGM from the explanation part

You can check the assigned range (slides X to X) for each track in the list.

📚 BGM Library

Uploaded BGM can be saved to the library and recalled from other projects with a single click.

💡 In sections where "Video audio only" is selected for video slides, the BGM will be automatically muted.


Chapter 13: OP/ED Branding

🎬 Inserting OP/ED Videos

Once you set your channel's intro and outro videos, they will be automatically combined with the main content during export. Supported formats are MP4 and WebM (under 500MB, no duration limit). Individual settings and saves are possible for each video size (16:9 / 9:16 / 1:1).

✂️ Trimming and Display Settings (v3.x)

  • Trimming: Specify the start and end seconds using the timeline handles (displayed as "Playing X to X seconds out of X seconds")

  • Display Method: Show All / Fill Screen / Stretch

  • Zoom Level

  • Margin Background: Black / Global Setting

You can check this on the spot in the video preview.

📚 OP/ED Library

Once uploaded, OP/ED videos are saved to the library and can be applied to future projects with a single click.


Chapter 14: Using Your Own API Settings (For Advanced Users)

This is for users who want to generate audio from the editor by setting their own API keys without using the Bundle Builder.

🔑 Obtaining a Gemini API Key

  1. Access Google AI Studio (aistudio.google.com)

  2. Log in with your Google account

  3. Click "Create API Key"

  4. Copy the issued key (a string starting with AIza)

  5. Paste it into the "Settings → Audio" tab of the app

You can perform the following operations on the settings screen.

  • Verify Key: Check validity, status, or insufficient permissions on the spot

  • Do not save to this device (session only): When ON, it will be cleared when the tab/browser is closed

  • Clear API Key: Delete the saved key from the browser

  • Paid Plan (Subscribed): When ON, the waiting time for batch generation is reduced (4 seconds → 1 second/item)

Gemini models (TTS supported) can be selected from the following.

  • Gemini 3.1 Flash TTS Preview (Latest)

  • Gemini 2.5 Flash Preview TTS (Fast)

  • Gemini 2.5 Pro Preview TTS (High Quality/For Paid Users)

⚠️ Pro TTS is likely to fail with free usage settings, so please use Flash TTS or enable the paid plan.

🔑 Obtaining a Google Cloud Text-to-Speech API Key (Optional)

Configure this if you want to use uniform and stable audio, although emotional expression is modest.

  1. Create a project in the Google Cloud Console

  2. Set up billing information

  3. Enable the Cloud Text-to-Speech API

  4. Create an API key and enter it into the app's settings screen


Chapter 15 Exporting (Videos, PDFs, and Various Materials)

🎬 Exporting Video Files (MP4/WebM)

When you press "Export Video," you can select the following on the confirmation screen.

  • Export Format: Auto / MP4 / WebM

  • Video Quality: Low (4Mbps) / Standard (8Mbps) / High (12Mbps)

  • Output Resolution: 720p / 1080p / 1440p

The frame rate is 30fps, and an estimated file size is also displayed.

ℹ️ MP4 may not be supported by some browsers; in such cases, it will automatically switch to WebM.


⚠️ Notes on Video Exporting

  • Do not close or reload the tab during processing.

  • Do not switch to other tabs, and pause any high-load tasks.

  • Connect laptops to a power source and avoid power-saving or sleep modes.

  • Check the available storage space on your destination drive.

  • If issues occur, try lowering the resolution or quality by one level and retry.

  • Reload the page after exporting (to free up memory).

If memory usage is high, a warning will be displayed before exporting.

⚠️ Regarding MP4 Compatibility

Even if you export as MP4, it may not be uploadable as-is. Since some social media platforms may not support certain standards, please convert to H.264 if necessary.

📄 Slide Output (PPTX / PDF / Images)

You can select the following formats from "Slide Output".

  • PPTX: Retain appearance as images and transcribe script to speaker notes

  • PDF (Slides only): Distribution PDF with 1 slide per page

  • PDF (Booklet): Combine slides and script on A4 portrait

  • Image ZIP (PNG): Image quality priority

  • Image ZIP (JPEG): File size priority

Additionally, you can configure the following settings:

  • Select by usage (Presets): Presentation sharing / Handouts / Image assets

  • Output content: Reflect edits / Original slides only

  • Pages to output: All / Current only / Range selection

  • File name

  • Output resolution

  • Output preview: Check the actual output image by paging through

  • Pre-export check: Warn in advance about missing assets, high-resolution load, ZIP size, etc.

📚 Detailed settings for Booklet PDF

This is a document PDF with "slide image + subtitle text" arranged on A4 portrait. If the subtitles are long, pages will be split automatically.

  • Subtitle Display Mode: Text/Speech Bubble

  • Character Name Display: Auto/Always Show/Hide

  • Embedded Slide Image Quality

Ideal for printed materials and seminar handouts.

ℹ️ Video slides export the first frame as a still image. The video itself is not included.


Chapter 16 Troubleshooting

🔇 No sound / Preview stops

Cause: The browser's "Autoplay Block" feature is active.
Solution: Click anywhere on the screen before playing. If it does not improve, reload the tab.

🐢 Performance is slow / Export errors occur

Cause: Insufficient browser memory (RAM). All video processing is performed on your computer.

  1. Delete unnecessary projects from the dashboard

  2. Close extra tabs or stop other tasks

  3. Lower the import quality or output resolution

  4. Restart the browser and try again

  5. Always reload the page after exporting

  6. For long videos, split the project and export individually

😱 Data is gone

Cause: IndexedDB is cleared by security software or browser "Clear browsing history/cookies" settings. It also disappears when closing Incognito mode.

Recovery is not possible from the app side. Regular backups (ZIP/JSON) are the only safeguard.

🖼️ Animated images cannot be loaded

The file size, frame count, playback duration, or total pixel count may have exceeded the limit. In v3.4, images that are likely to exceed the limit are automatically downscaled where possible, and the post-downscaling size will be notified. If an image cannot be loaded, a message explaining the cause will be displayed.

🎭 Expressions and gestures are not reflected

Expressions and gestures are not reflected in styles that do not display character images (such as banners). Also, please check if the assets required for performance (MouthLoop character pack) are registered.

📊 Displayed time or file size differs from reality

The statistical information in the header is an estimate. It may differ slightly from the actual exported file.

🔑 Error occurs during voice generation

  • "API key not set" → Enter the key in Settings → Voice tab

  • "Rate limit" → Wait a while and try again. Please also note the daily limit for the free tier

  • "API key rejected" → Check if a Gemini key was entered in the Google Cloud field or check the key restriction settings

  • "Google Cloud API disabled" → Enable the Cloud Text-to-Speech API in the Console


About updating

If you have already purchased SlideCast Studio, please copy the latest version from the same paid note shared link you have been using.

If you wish to carry over existing projects, libraries, or settings, please create a backup from "Save Data" first just in case.

To update the GAS version while keeping the same URL, replace the following files in your existing GAS project with the latest version files:

  • Untitled.html

  • AppBundle.html

Then, redeploy the existing deployment as a "New Version".

Please see here for detailed steps.

If you are using the local HTML version, please download and open "SlideCast Studio [Local Version].html".

If you open it in a new GAS project or a different browser, previous data will not be displayed automatically. In that case, please restore from a backup.


Conclusion

SlideCast Studio is continuously being updated. Please follow this note and the creator's X account to check for the latest information. If you have any questions, requests, or bug reports, please leave them in the comments.


いいなと思ったら応援しよう!

aokuma 創作活動の継続のために よろしければ応援お願いします! いただいたチップはクリエイターとしての活動費に使わせていただきます!