Skip to content

Real-Time Streaming ​

Last Updated: September 2026

Scrapalot streams AI responses in real time, so you see answers as they're generated, together with status updates, citations and research progress.

How Streaming Works ​

When you ask a question, Scrapalot:

  1. Immediately sends status updates (which strategy was chosen, what it is doing)
  2. Streams the answer token by token as it's generated
  3. Sends citations and other rich content (charts, research plans, suggestions) as separate packets
  4. Ends with a stream_end packet that says why the stream finished

Why streaming matters:

  • No waiting for the full response
  • See progress as it happens
  • Stop the answer if it is going in the wrong direction (the partial answer is kept)

What You See ​

Status Updates ​

While a question is being processed, the chat shows short, translated status lines such as "Retrieving documents", "Reranking documents" or "Expanding context". The server sends these as status codes (for example retrieving_documents, reranking_documents, expanding_context), not as English text; the app translates them into your interface language.

Answer Streaming ​

The answer appears word by word, rendered as Markdown with inline citation markers. You can start reading before it is complete and stop it at any time.

Citations ​

Each source arrives as its own citation_info packet with the document title, page, relevance score, a short excerpt and (for PDFs) the position used to highlight the passage in the viewer.

Completion ​

The stream ends with stream_end, which carries the reason plus usage details when available: total/input/output tokens, duration, tokens per second, cost, provider and model.

Connection Types ​

Server-Sent Events (chat answers) ​

Chat answers — in the web app and for API clients — are streamed over Server-Sent Events from one endpoint:

POST /api/v1/chat/completions

Send "stream": true for an SSE stream of chat.completion.chunk objects, or omit it for a single chat.completion JSON response. The endpoint is wire-compatible with the OpenAI API, so standard openai SDK clients work when pointed at the /api/v1 base. See the API Reference for request fields.

Inside the SSE stream:

  • Answer tokens (message_delta packets) arrive as choices[0].delta.content.
  • Every other packet is relayed verbatim as choices[0].delta.scrapalot — an object with a type field and the packet's own fields (see the catalogue below).
  • The stream finishes with a chunk carrying finish_reason: "stop" and then data: [DONE].
data: {"id":"chatcmpl-…","object":"chat.completion.chunk",…,"choices":[{"index":0,"delta":{"scrapalot":{"type":"status","content":"retrieving_documents","stage":null}}}]}

data: {"id":"chatcmpl-…","object":"chat.completion.chunk",…,"choices":[{"index":0,"delta":{"content":"According to"}}]}

data: {"id":"chatcmpl-…","object":"chat.completion.chunk",…,"choices":[{"index":0,"delta":{"scrapalot":{"type":"stream_end","reason":"completed","total_tokens":812}}}]}

data: {"id":"chatcmpl-…","object":"chat.completion.chunk",…,"choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]

A long-running deep research run can be re-attached after a disconnect with GET /api/v1/research/{plan_id}/stream?from_id=<last id>, an SSE stream whose events carry a resumable id:.

WebSocket (STOMP) ​

The web app also keeps a STOMP-over-WebSocket connection open through the gateway. It is used for background events — document-processing job progress, notifications, workspace chat and collaboration — not for chat answers. Job progress on this channel is a percentage (0–100). Details: WebSocket Architecture.

Packet Format ​

Internally every packet is one line of JSON with a sequence index and the packet object:

json
{"ind": 3, "obj": {"type": "citation_info", "timestamp": "2026-09-18T10:00:00Z", "citation_num": 1, "document_id": "…", "document_title": "…", "page": 5, "score": 0.92}}
  • ind — sequence number, increasing within a stream
  • obj.type — the packet type (below)
  • obj.timestamp — UTC time the packet was created
  • the remaining fields depend on the type

Over the OpenAI-compatible SSE stream you receive obj as delta.scrapalot (or, for message_delta, just its text as delta.content).

Packet Catalogue ​

Scrapalot defines 94 packet types. Clients should ignore types they do not recognise — new ones are added as features ship.

GroupPacket types
Message (2)message_start, message_delta
Reasoning (2)reasoning_start, reasoning_delta
Model insight (2)model_insight_start, model_insight_delta — the model's own-knowledge reflection when "Thinking" is on
Citations (5)citation_start, citation_info, citation_delta, strategy_transparency, graph_expansion
Status & control (9)status, error, collection_scope_expanded, transcription_partial, audio_delta, transcription_final, image_attached, stream_end, section_end
Tools (3)tool_start, tool_delta, tool_artifact
Research (7)research_start, research_query, research_result, research_plan, planning_progress, research_section, research_thinking
Task execution (5)task_decomposition_plan, task_execution, parallel_group, quality_gate, task_coordination
Search (6)search_strategy, source_evaluation, search_progress, content_extraction, search_fusion, result_ranking
Multi-agent coordination (5)coordination_plan, agent_status, inter_agent_communication, coordination_quality_gate, agent_pool_status
Synthesis & quality (8)synthesis_start, synthesis_delta, validation_start, validation_result, quality_check, quality_result, citation, report_generation
Context expansion (3)context_expansion, chapter_read_start, chapter_read_progress
Intent routing (1)intent_routing
Clarification & report (6)clarification_questions, plan_preview, research_report, discovery_start, discovery, discovery_complete
Paper generation (3)paper_progress, paper_complete, paper_error
Iterative research (6)iteration_start, iteration_complete, reflection_complete, continuation_decision, hypothesis, iteration_state
Council (3)council_start, council_member, council_synthesis
Research v2 (5)agent_persona, source_curation, research_cost, review_feedback, revision
Multi-step research (3)research_step_start, research_step_complete, research_gap_analysis
Debug (1)rag_debug_info
Suggestions & charts (2)suggestions, chart_data
Knowledge-graph communities (4)community_build_started, community_build_progress, community_report_generated, community_build_complete
Multimodal processing (3)multimodal_element_started, multimodal_element_described, multimodal_element_indexed

Packets most clients need ​

TypeKey fieldsMeaning
message_deltacontentA piece of the answer text
reasoning_deltacontentA piece of the model's visible reasoning
statuscontent (status code), stageProcessing step, e.g. retrieving_documents
citation_infocitation_num, document_id, document_title, page, score, text, citation_context, stanceOne cited source
strategy_transparencystrategy detailsWhich retrieval strategy ran and why
chart_datachart_type (bar, line, pie, scatter, area, radar, composed, heatmap, boxplot), dataA chart to render inline
suggestionssuggestion listFollow-up questions
clarification_questionsquestions, request_idDeep research asks before planning
plan_previewplan_id, title, sections, total_questions, estimated_duration_minutesResearch plan waiting for approval
research_reportplan_id, title, executive_summary, full_report_markdown, quality_score, total_sourcesFinal deep-research report
errorcontent, error_codeSomething failed; followed by stream_end with reason: "error"
stream_endreason, token/cost/latency fields, provider, modelEnd of stream

stream_end.reason is one of completed, error, cancelled, clarification_needed, plan_preview_ready or background_job_dispatched (a research run that continues as a background job).

User Experience ​

Phase 1: Query analysis — status updates and, if enabled, the strategy selection shown in the transparency panel.

Phase 2: Retrieval — status updates while documents are searched, reranked and expanded.

Phase 3: Answer generation — streaming text with citations.

Phase 4: Complete — usage details and follow-up suggestions.

Cancellation ​

Click the stop button (or call POST /api/v1/chat/sessions/{session_id}/stop). Generation stops and the partial answer is saved to the conversation.

Performance ​

Typical behaviour:

  • Status updates appear almost immediately
  • First answer token usually within a few seconds (depends on model and strategy)
  • Deep research runs for minutes, not seconds

Factors affecting speed: the AI model, the retrieval strategy, the size of the collections searched, query complexity and server load.

Troubleshooting ​

No streaming, only the full response ​

  • Make sure the request sets "stream": true and Accept: text/event-stream
  • With curl, use --no-buffer
  • A reverse proxy in front of a self-hosted install must not buffer responses (for Nginx, proxy_buffering off on the API location)

Stream stops early ​

  • Check the last stream_end.reason and any preceding error packet
  • Proxy or load-balancer read timeouts must allow long responses — deep research can stream for many minutes

WebSocket keeps reconnecting ​

  • Proxies must pass the Upgrade and Connection headers
  • Idle timeouts on the proxy should be longer than the STOMP heartbeat

Open-core — Community Edition under AGPL-3.0 · Hosted product is proprietary.