## ADDED Requirements ### Requirement: Voice call mode When `features.realtimeVoiceCall` is enabled and a usable `realtimeVoice` endpoint is resolved, the phone control on the bottom toolbar SHALL start an in-app voice call session with the AI (without requiring the Ask modal). Course media audio MUST pause for the call duration. The session MUST support mute and hang up, and on hang up return to idle with course media remaining paused. #### Scenario: Start and end voice call - **WHEN** the learner taps the phone control with a valid realtime voice configuration - **THEN** a voice call session starts, course playback is paused, and hanging up ends the session without auto-resuming course media ### Requirement: Video toolbar opens preview; screen-share from modal The bottom-toolbar video control SHALL open the Ask modal with video preview rather than immediately starting a screen-share call. When `features.realtimeScreenShare` is enabled and a usable realtime vision (or equivalent) endpoint is resolved, the Ask modal MUST expose a screen-share / 投屏通话 action that starts an in-app full-screen share session combining realtime voice with sparse visual frames from the study stage (not system desktop capture). #### Scenario: Video icon opens preview first - **WHEN** the learner taps the video toolbar icon - **THEN** the Ask modal opens with video preview and does not auto-start screen-share #### Scenario: Start screen-share from modal - **WHEN** the learner activates 投屏通话 inside the video Ask modal with a valid configuration - **THEN** the app enters an in-app full-screen share UI, streams microphone audio, and uploads study-stage JPEG frames at a sparse cadence (~1 fps or speech/event-driven) #### Scenario: Hang up screen-share call - **WHEN** the learner hangs up a screen-share call - **THEN** frame upload and voice session stop, the full-screen share UI dismisses, and course media remains paused ### Requirement: Sparse frame default path Default visual cadence for screen-share MUST be approximately 1 frame per second or event-driven, and MUST NOT stream continuous high-FPS video to the model as the default path. #### Scenario: Cost-controlled framing - **WHEN** a screen-share call is active - **THEN** visual uploads remain at sparse cadence suitable for cost control ### Requirement: Call mode mutual exclusion Annotate mode, voice call, and screen-share call MUST be mutually exclusive. Entering one MUST end or block the others. While a call is active, annotate and the opposite call control MUST be disabled or end the current call before switching (implementation MAY choose disable-only in v1). #### Scenario: Cannot annotate during call - **WHEN** a voice or screen-share call is active - **THEN** activating the pencil tool does not enter annotate over the live call without ending the call first (disabled or explicit end-then-annotate) ### Requirement: Optional RTC provider upgrade path The system MAY offer a provider RTC / conversational-AI path (Tier B) when `rtcAppId` and token configuration are present, but the default product path MUST remain the sparse-frame Tier A described above. Absence of RTC configuration MUST NOT block Tier A. #### Scenario: Tier A works without RTC app id - **WHEN** `rtcAppId` is empty but realtime voice/vision endpoints are configured - **THEN** voice and screen-share call actions still operate via the sparse-frame / realtime WebSocket (or equivalent) Tier A path