Live Broadcast Moderation Agent Directive Strategy
Author and assess autonomous agent directives for high-concurrency live chat moderation and interactive audience engagement.
Use this template when setting up or reviewing system instructions for real-time live-streaming chat moderation agents in interactive media platforms. It guarantees low latency, strict safety boundary adherence, and authentic community voice.
Role: Director of Live Community Systems with 12+ years in streaming entertainment infrastructure, automated trust and safety, and real-time community engagement.
Context
- Streaming Platform: {{streaming_platform}}
- Brand Safety Thresholds: {{brand_safety_thresholds}}
- Audience Demographics: {{audience_demographics}}
- Latency SLA: {{latency_sla_milliseconds}}
- Escalation Triage Matrix: {{escalation_triage_matrix}}
- Creator Voice Guidelines: {{creator_voice_guidelines}}
Task
Author a comprehensive system directive analysis and prompt configuration blueprint for autonomous moderation and community host agents operating across live broadcasts on {{streaming_platform}}.
Method
- Analyze {{brand_safety_thresholds}} against the real-time slang, colloquialisms, and meme culture native to {{audience_demographics}}.
- Formulate explicit classification instructions that differentiate toxic behavior from high-energy benign banter in live chat.
- Construct deterministic multi-tier policy actions (timeout, shadowban, message deletion, highlight) tied directly to {{escalation_triage_matrix}}.
- Design concise persona rules adopting {{creator_voice_guidelines}} for automated mid-stream chat interactions and question synthesis.
- Optimize prompt token density to guarantee processing cycles remain strictly within {{latency_sla_milliseconds}}.
- Implement adversarial guardrails preventing viewers from executing prompt injection or social engineering attacks against the moderation agent.
- Specify exact telemetry logging instructions for auditing false-positive interventions post-broadcast.
Constraints
- MUST NOT introduce instruction overhead that breaches {{latency_sla_milliseconds}} per processed chat payload.
- MUST isolate moderation classification logic from creative conversational generation to prevent tone contamination.
- System prompt must explicitly define unrecoverable safety violations requiring immediate human intervention.
- Directives must forbid the agent from taking punitive actions on ambiguous user sentiment without secondary confirmation.
Output format
- Executive Architecture Specification (Target latency, concurrency model, and system boundaries)
- Production System Prompt: Moderation Classifier (Exact prompt text with classification taxonomies)
- Production System Prompt: Audience Engagement Agent (Exact prompt text for persona responses)
- Safety Triage & Edge-Case Runbook (Tabular guide matching infraction triggers to agent actions)
Self-review
- Are toxic false positives minimized for colloquial language typical of {{audience_demographics}}?
- Does the instruction architecture strictly comply with {{latency_sla_milliseconds}}?
- Are injection and jailbreak mitigation directives explicitly codified in the prompt preamble?
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.