Real-Time Streaming
Last Updated: September 2026
Scrapalot streams AI responses in real time, so you see answers as they're generated, together with status updates, citations and research progress.
How Streaming Works
When you ask a question, Scrapalot:
- Immediately sends status updates (which strategy was chosen, what it is doing)
- Streams the answer token by token as it's generated
- Sends citations and other rich content (charts, research plans, suggestions) as separate packets
- Ends with a
stream_endpacket that says why the stream finished
Why streaming matters:
- No waiting for the full response
- See progress as it happens
- Stop the answer if it is going in the wrong direction (the partial answer is kept)
What You See
Status Updates
While a question is being processed, the chat shows short, translated status lines such as "Retrieving documents", "Reranking documents" or "Expanding context". The server sends these as status codes (for example retrieving_documents, reranking_documents, expanding_context), not as English text; the app translates them into your interface language.
Answer Streaming
The answer appears word by word, rendered as Markdown with inline citation markers. You can start reading before it is complete and stop it at any time.
Citations
Each source arrives as its own citation_info packet with the document title, page, relevance score, a short excerpt and (for PDFs) the position used to highlight the passage in the viewer.
Completion
The stream ends with stream_end, which carries the reason plus usage details when available: total/input/output tokens, duration, tokens per second, cost, provider and model.
Connection Types
Server-Sent Events (chat answers)
Chat answers — in the web app and for API clients — are streamed over Server-Sent Events from one endpoint:
POST /api/v1/chat/completionsSend "stream": true for an SSE stream of chat.completion.chunk objects, or omit it for a single chat.completion JSON response. The endpoint is wire-compatible with the OpenAI API, so standard openai SDK clients work when pointed at the /api/v1 base. See the API Reference for request fields.
Inside the SSE stream:
- Answer tokens (
message_deltapackets) arrive aschoices[0].delta.content. - Every other packet is relayed verbatim as
choices[0].delta.scrapalot— an object with atypefield and the packet's own fields (see the catalogue below). - The stream finishes with a chunk carrying
finish_reason: "stop"and thendata: [DONE].
data: {"id":"chatcmpl-…","object":"chat.completion.chunk",…,"choices":[{"index":0,"delta":{"scrapalot":{"type":"status","content":"retrieving_documents","stage":null}}}]}
data: {"id":"chatcmpl-…","object":"chat.completion.chunk",…,"choices":[{"index":0,"delta":{"content":"According to"}}]}
data: {"id":"chatcmpl-…","object":"chat.completion.chunk",…,"choices":[{"index":0,"delta":{"scrapalot":{"type":"stream_end","reason":"completed","total_tokens":812}}}]}
data: {"id":"chatcmpl-…","object":"chat.completion.chunk",…,"choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]A long-running deep research run can be re-attached after a disconnect with GET /api/v1/research/{plan_id}/stream?from_id=<last id>, an SSE stream whose events carry a resumable id:.
WebSocket (STOMP)
The web app also keeps a STOMP-over-WebSocket connection open through the gateway. It is used for background events — document-processing job progress, notifications, workspace chat and collaboration — not for chat answers. Job progress on this channel is a percentage (0–100). Details: WebSocket Architecture.
Packet Format
Internally every packet is one line of JSON with a sequence index and the packet object:
{"ind": 3, "obj": {"type": "citation_info", "timestamp": "2026-09-18T10:00:00Z", "citation_num": 1, "document_id": "…", "document_title": "…", "page": 5, "score": 0.92}}ind— sequence number, increasing within a streamobj.type— the packet type (below)obj.timestamp— UTC time the packet was created- the remaining fields depend on the type
Over the OpenAI-compatible SSE stream you receive obj as delta.scrapalot (or, for message_delta, just its text as delta.content).
Packet Catalogue
Scrapalot defines 94 packet types. Clients should ignore types they do not recognise — new ones are added as features ship.
| Group | Packet types |
|---|---|
| Message (2) | message_start, message_delta |
| Reasoning (2) | reasoning_start, reasoning_delta |
| Model insight (2) | model_insight_start, model_insight_delta — the model's own-knowledge reflection when "Thinking" is on |
| Citations (5) | citation_start, citation_info, citation_delta, strategy_transparency, graph_expansion |
| Status & control (9) | status, error, collection_scope_expanded, transcription_partial, audio_delta, transcription_final, image_attached, stream_end, section_end |
| Tools (3) | tool_start, tool_delta, tool_artifact |
| Research (7) | research_start, research_query, research_result, research_plan, planning_progress, research_section, research_thinking |
| Task execution (5) | task_decomposition_plan, task_execution, parallel_group, quality_gate, task_coordination |
| Search (6) | search_strategy, source_evaluation, search_progress, content_extraction, search_fusion, result_ranking |
| Multi-agent coordination (5) | coordination_plan, agent_status, inter_agent_communication, coordination_quality_gate, agent_pool_status |
| Synthesis & quality (8) | synthesis_start, synthesis_delta, validation_start, validation_result, quality_check, quality_result, citation, report_generation |
| Context expansion (3) | context_expansion, chapter_read_start, chapter_read_progress |
| Intent routing (1) | intent_routing |
| Clarification & report (6) | clarification_questions, plan_preview, research_report, discovery_start, discovery, discovery_complete |
| Paper generation (3) | paper_progress, paper_complete, paper_error |
| Iterative research (6) | iteration_start, iteration_complete, reflection_complete, continuation_decision, hypothesis, iteration_state |
| Council (3) | council_start, council_member, council_synthesis |
| Research v2 (5) | agent_persona, source_curation, research_cost, review_feedback, revision |
| Multi-step research (3) | research_step_start, research_step_complete, research_gap_analysis |
| Debug (1) | rag_debug_info |
| Suggestions & charts (2) | suggestions, chart_data |
| Knowledge-graph communities (4) | community_build_started, community_build_progress, community_report_generated, community_build_complete |
| Multimodal processing (3) | multimodal_element_started, multimodal_element_described, multimodal_element_indexed |
Packets most clients need
| Type | Key fields | Meaning |
|---|---|---|
message_delta | content | A piece of the answer text |
reasoning_delta | content | A piece of the model's visible reasoning |
status | content (status code), stage | Processing step, e.g. retrieving_documents |
citation_info | citation_num, document_id, document_title, page, score, text, citation_context, stance | One cited source |
strategy_transparency | strategy details | Which retrieval strategy ran and why |
chart_data | chart_type (bar, line, pie, scatter, area, radar, composed, heatmap, boxplot), data | A chart to render inline |
suggestions | suggestion list | Follow-up questions |
clarification_questions | questions, request_id | Deep research asks before planning |
plan_preview | plan_id, title, sections, total_questions, estimated_duration_minutes | Research plan waiting for approval |
research_report | plan_id, title, executive_summary, full_report_markdown, quality_score, total_sources | Final deep-research report |
error | content, error_code | Something failed; followed by stream_end with reason: "error" |
stream_end | reason, token/cost/latency fields, provider, model | End of stream |
stream_end.reason is one of completed, error, cancelled, clarification_needed, plan_preview_ready or background_job_dispatched (a research run that continues as a background job).
User Experience
Phase 1: Query analysis — status updates and, if enabled, the strategy selection shown in the transparency panel.
Phase 2: Retrieval — status updates while documents are searched, reranked and expanded.
Phase 3: Answer generation — streaming text with citations.
Phase 4: Complete — usage details and follow-up suggestions.
Cancellation
Click the stop button (or call POST /api/v1/chat/sessions/{session_id}/stop). Generation stops and the partial answer is saved to the conversation.
Performance
Typical behaviour:
- Status updates appear almost immediately
- First answer token usually within a few seconds (depends on model and strategy)
- Deep research runs for minutes, not seconds
Factors affecting speed: the AI model, the retrieval strategy, the size of the collections searched, query complexity and server load.
Troubleshooting
No streaming, only the full response
- Make sure the request sets
"stream": trueandAccept: text/event-stream - With curl, use
--no-buffer - A reverse proxy in front of a self-hosted install must not buffer responses (for Nginx,
proxy_buffering offon the API location)
Stream stops early
- Check the last
stream_end.reasonand any precedingerrorpacket - Proxy or load-balancer read timeouts must allow long responses — deep research can stream for many minutes
WebSocket keeps reconnecting
- Proxies must pass the
UpgradeandConnectionheaders - Idle timeouts on the proxy should be longer than the STOMP heartbeat
Related Documentation
- API Reference - API endpoints
- WebSocket Architecture - STOMP channels
- RAG Strategy - How queries are processed
- Deep Research - Research phases and packets
- User Guide - Using the chat interface