# Spike notes (tasks 4.4 & 6.4) ## 4.4 Audio clip upload vs ASR (v1 pick) **Decision for v1:** time-range + mid-frame vision (and prompt text). Do **not** require raw audio upload yet. | Path | Pros | Cons | |------|------|------| | Audio bytes to multimodal | Faithful dialogue | Ark/OpenAI chat endpoints differ; many need separate ASR/file APIs | | Cloud ASR → text | Works with chat-only keys | Extra latency/cost; speaker diarization weak | | **Time range + mid-frame (v1)** | Uses existing vision/chat; ships now | Model infers dialogue from frame + timestamps | Follow-up: when `realtimeVoice` / ASR endpoint is confirmed for the active profile, attach transcript of `[start,end]` before send. ## 6.4 Tier B RTC (Phase 2) Tier A (default) does **not** need `rtcAppId` / `rtcTokenUrl`. **Phase 2 (not blocking):** Volcengine conversational AI / RTC (`StartVoiceChat`) when both `rtcAppId` and token URL are configured — closer to Doubao in-app video call. Document only until product enables RTC keys for StudyDeck.