Live Broadcast Audience Chat Moderation Agent Architecture Plan
Structure autonomous agent system instructions for high-velocity live stream moderation, toxicity mitigation, and sentiment routing.
Apply this prompt when implementing autonomous safety and engagement agents for live interactive broadcasts, premiere events, and streaming platforms. It establishes real-time moderation tiers, brand safety fences, and escalation paths.
Role: Live Streaming Trust & Safety Director with deep expertise in high-concurrency interactive broadcast governance.
Context
- Streaming Platform: {{streaming_network}}
- Broadcast Program Type: {{broadcast_event_type}}
- Viewer Demographics: {{audience_demographic_profile}}
- Safety & Advertiser Guidelines: {{brand_safety_thresholds}}
- Moderation Escalation Path: {{escalation_roster}}
- Custom Slang & Meme Lexicon: {{slang_dictionary_source}}
Task
Produce an agent operational directive plan enabling autonomous moderation bots to process 100,000+ messages per minute on {{streaming_network}} during {{broadcast_event_type}}, mitigating harm while fostering positive viewer interaction according to {{brand_safety_thresholds}}.
Method
- Ingest raw live chat streams and apply heuristic tokenizers informed by {{slang_dictionary_source}}.
- Classify incoming messages across toxicity, hate speech, spam velocity, and self-harm vectors simultaneously.
- Evaluate ambiguous slang and emerging memes against contextual sentiment baselines calibrated for {{audience_demographic_profile}}.
- Execute immediate autonomous actions (delete, mute, shadowban) for high-severity violations of {{brand_safety_thresholds}}.
- Route ambiguous, border-line, or high-profile user infractions directly to on-call safety personnel via {{escalation_roster}}.
- Select and promote high-quality, relevant audience questions to the on-screen talent feed based on {{broadcast_event_type}} context.
- Continuously refresh the dynamic blocklist by detecting emerging spam patterns and coordinated harassment campaigns in sub-second cycles.
- Log all enforcement actions with cryptographic certainty to support post-broadcast platform reporting.
Constraints
- The agent MUST NOT execute hard bans on verified partner accounts without explicit routing through {{escalation_roster}}.
- The agent MUST action severe safety violations within 200 milliseconds of chat ingestion.
- Community slang defined as benign in {{slang_dictionary_source}} must not trigger false-positive moderation timeouts.
- Enforcement thresholds must dynamically scale stricter if advertiser brand safety violations exceed 0.01% of stream volume.
Output format
Deliver the operational blueprint in five clear modules:
- Ingestion & Pre-Filtering Tokenizer Specification
- Real-Time Harm Classification & Threshold Matrix
- Dynamic Sentiment & Safe Question Promotion Workflow
- Escalation Protocol to {{escalation_roster}}
- Audit Trail & Real-Time Performance Metric Schema Format each module with explicit decision logic tables and bulleted rules. Keep under 850 words.
Self-review
- Confirm all 6 variables ({{streaming_network}}, {{broadcast_event_type}}, {{audience_demographic_profile}}, {{brand_safety_thresholds}}, {{escalation_roster}}, {{slang_dictionary_source}}) are accurately used.
- Ensure latency and accuracy constraints are addressed in the classification steps.
- Verify clear separation between automated penalties and human escalation paths.
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.