## ADDED Requirements ### Requirement: Previous-sentence energy boundary clip When the learner activates previous-sentence on a paused (or pause-on-activate) media position `t`, the system SHALL locate a low-energy boundary by searching `[t-5s, t)` first, then `[t-10s, t)` if needed, using short-window energy (e.g. RMS) with a minimum silence duration. If no valley is found, the system MUST fall back to `max(0, t-10s)`. The resulting interval `[start, t]` SHALL be treated as the ask clip (highlighted on the waveform when shown) and SHALL open the Ask modal with an audio-clip preview bounded to that range. #### Scenario: Valley found within 5 seconds - **WHEN** the learner taps previous-sentence and a qualifying low-energy valley exists within the prior 5 seconds - **THEN** the ask clip start equals that valley time and the end equals the current media time #### Scenario: Expand search to 10 seconds then hard fallback - **WHEN** no qualifying valley exists in the prior 5 seconds - **THEN** the system searches the prior 10 seconds, and if still none, uses `max(0, t-10s)` as the clip start ### Requirement: Previous-sentence clip ask to model Submitting a previous-sentence ask SHALL send the clip time range (and audio bytes and/or ASR text when the configured endpoint supports them) together with the learner question and the video/audio ask prompt, so the model can answer questions such as segment meaning, speaker content, or relation to prior context. #### Scenario: Ask about dialogue meaning - **WHEN** the learner confirms a previous-sentence clip and sends a question about what was said - **THEN** the model request includes the clip time range (and available audio or transcript context) plus the configured ask prompt ### Requirement: Ask modal with preview and conversation panes The system SHALL present ask interactions in a modal overlay with: - **Left**: media preview for the active context — image (screenshot/region), audio clip (bounded play), or video (seekable preview); image preview SHOULD support zoom and download where practical. - **Right**: conversation — prior user/AI turns, streaming or incremental assistant text rendered as Markdown when applicable, composer (text + voice input), and send. Esc or an explicit close control MUST dismiss the modal without auto-resuming course media. #### Scenario: Open from screenshot chip - **WHEN** the learner taps a screenshot thumbnail chip on the toolbar - **THEN** the Ask modal opens with that image in the left preview and the conversation pane ready for input #### Scenario: Close keeps media paused - **WHEN** the learner closes the Ask modal after a pause-for-ask flow - **THEN** course media remains paused ### Requirement: Message entry pauses and captures a full frame Activating the message control (or equivalent composer entry that starts an image ask) SHALL pause playback, capture a full-frame still into the capture buffer, and open the Ask modal with image preview. #### Scenario: Message while playing - **WHEN** video is playing and the learner taps the message control - **THEN** playback pauses, a full-frame screenshot is stored, and the Ask modal opens with that image ### Requirement: Pencil annotate with mask and shape tools Activating the pencil tool SHALL pause playback, apply a dark semi-transparent full-stage mask, and present a QQ-style capture toolbar with at least rectangle, ellipse, confirm, and cancel. The selected region MUST punch through the mask so original colors remain fully visible inside the selection. Confirming SHALL crop the selection into the capture buffer, exit annotate mode, keep playback paused, and open the Ask modal with the cropped image preview. #### Scenario: Confirm elliptical region - **WHEN** the learner draws an ellipse selection and taps confirm - **THEN** the cropped region image is placed in the capture buffer, the mask is dismissed, media remains paused, and the Ask modal opens on that image #### Scenario: Cancel annotate - **WHEN** the learner cancels annotate mode - **THEN** the mask and tools dismiss without writing a new buffer item and without opening the Ask modal ### Requirement: Capture buffer for vision ask Sending from the Ask modal SHALL attach the capture buffer image when present. Requests with an image MUST use a vision-capable endpoint and the video/image ask prompt. Successful send MUST clear or advance buffer state per UX rules; failed send MUST retain the buffer for retry. #### Scenario: Send with buffered screenshot - **WHEN** a capture buffer image exists and the learner sends a non-empty question from the Ask modal - **THEN** the client calls a vision endpoint with the image and question ### Requirement: TTS playback for assistant messages Each assistant text message in the Ask modal SHALL expose a speaker control. Playback MUST prefer edgeTTS (local CLI / Tauri-invoked `edge-tts`). If edgeTTS is unavailable, quota-limited, or fails, the system MUST fall back to on-device `speechSynthesis` (iPad/local speaker) and MAY show a short non-blocking hint. #### Scenario: edgeTTS success - **WHEN** the learner taps the speaker on an assistant message and edgeTTS is available - **THEN** the message is spoken via edgeTTS audio #### Scenario: Fallback to local speech - **WHEN** edgeTTS fails or is unavailable - **THEN** the same message is spoken via `speechSynthesis` without blocking further chat ### Requirement: WeChat-style voice input messages The Ask modal composer SHALL provide a speaker/mic control that toggles recording: first tap starts recording, second tap stops and sends a voice message bubble. Voice bubbles MUST be replayable on tap. The system SHOULD attach speech-to-text transcript when the platform provides it and include transcript (and current visual context when relevant) in the model request; if transcript is missing, the voice bubble is still sent and the ask MAY proceed with visual/time context plus a generic voice-ask prompt. #### Scenario: Record and send voice bubble - **WHEN** the learner taps the voice-input control, speaks, then taps again to finish - **THEN** a user voice bubble appears in the conversation, is playable, and a model request is initiated using transcript and/or contextual media #### Scenario: Replay voice bubble - **WHEN** the learner taps a sent voice bubble - **THEN** that recording plays (and stops any other in-modal voice/TTS playback)