## 1. Stage chrome shell - [x] 1.1 Replace native video `controls` with a `StudyBottomChrome` layout (stage flex + wave strip + toolbar) on PC/iPad video modes in `StudyView` / `MediaPlayer` - [x] 1.2 Implement waveform strip: progress, time label, play/pause, seek; reserve ~56–64px (+ safe-area on iPad) - [x] 1.3 Implement toolbar scaffold: previous-sentence, phone, video, pencil, message icon, composer, send; gate by `AiFeatures` / endpoints - [x] 1.4 Remove or hide `ai-fab` as primary entry on video stages; keep answer bubble above toolbar ## 2. Capture buffer and message ask - [x] 2.1 Add single-slot `CaptureBuffer` (dataUrl, mediaTime, kind) with thumbnail chip + clear - [x] 2.2 On composer focus/first tap: pause video and full-frame capture into buffer - [x] 2.3 Wire send path: buffer present → vision + `videoAsk`; success clears buffer; failure retains - [x] 2.4 Show lightweight Q/A bubble above toolbar for replies ## 3. Pencil annotate overlay - [x] 3.1 Build `AnnotateOverlay`: pause, freeze frame canvas, dark mask, punch-hole selection - [x] 3.2 Add QQ-style tools: cancel, rectangle, ellipse, confirm; allow resize after draw - [x] 3.3 On confirm: crop selection into `CaptureBuffer`, exit annotate, keep paused - [x] 3.4 Enforce mutual exclusion with call modes (disable pencil while calling) ## 4. Previous-sentence energy clip - [x] 4.1 Implement energy series extraction (RMS windows) for current media URL with simple cache - [x] 4.2 Implement `EnergySentenceFinder`: search 5s then 10s valleys + hard fallback `t-10s` - [x] 4.3 Highlight `[start, t]` on waveform; previous-sentence button sets clip ready state - [x] 4.4 Spike: audio bytes upload vs ASR-to-text for configured Ark/OpenAI endpoints; pick one for v1 - [x] 4.5 Send clip ask (time range + audio/ASR + question + prompt); support meaning / speakers / context questions ## 5. Realtime voice call (Tier A) - [x] 5.1 Add `CallSession` voice path using `realtimeVoice` endpoint; request mic permission with clear errors - [x] 5.2 Pause course media on call start; provide mute + hang up; remain paused after hang up - [x] 5.3 Apply `prompts.voiceCall`; respect `features.realtimeVoiceCall` ## 6. Screen-share video call (Tier A sparse frames) - [x] 6.1 Add in-app full-screen share UI mirroring study stage (not system display capture) - [x] 6.2 Stream mic audio + sparse JPEG frames (~1 fps or speech-driven) via `realtimeVision` / multimodal realtime adapter - [x] 6.3 Hang up stops audio+frames and dismisses share UI; gate with `features.realtimeScreenShare` - [x] 6.4 Document Tier B (RTC `rtcAppId` / conversational AI) as Phase 2; ensure empty RTC config does not block Tier A ## 7. Polish and verification - [x] 7.1 iPad safe-area / touch targets; PC keyboard focus and Esc cancel annotate - [x] 7.2 Dual-video / video_pdf layout smoke: chrome does not clip PDF pane incorrectly - [x] 7.3 Manual checklist: message focus capture, pencil crop ask, previous-sentence clip, voice call, sparse-frame share call ## 8. Ask modal + voice UX (from polish chat) - [x] 8.1 Replace toolbar emoji with SVG icons (`IconPhone` / `IconVideo` / `IconPencil` / `IconMessage`) - [x] 8.2 Add `StudyAskModal` split preview (image/audio/video) + conversation; wire toolbar/chip entry points - [x] 8.3 Route screen-share start through video modal action (video icon opens preview first) - [x] 8.4 Assistant message TTS via edgeTTS with `speechSynthesis` fallback - [x] 8.5 WeChat-style voice input: tap-to-record, voice bubble, tap-to-play - [ ] 8.6 Fold user-reported interaction bugs into `design.md` UX backlog, then promote each to specs + fix - [ ] 8.7 Re-verify modal ↔ toolbar state machine (no dual composers / stuck overlays) ## 9. UX backlog intake (ongoing) - [ ] 9.1 Collect fragmented issues from hand-testing (paste into `design.md` → UX backlog) - [ ] 9.2 For each accepted issue: add/adjust requirement scenario in specs, then implement