Skip to content

Guardrails

The agent has limits built in so it stays a helpful game expert and nothing else. Some guardrails are fixed, and some you control. This page covers both.

On topic

The agent answers questions about your game. It won’t be pulled into unrelated debates, it won’t write a player’s essay, and it won’t role-play as something it isn’t. If a message isn’t about the game, the agent stays out of it, even in a channel where auto mode is on.

What it refuses

Some things the agent won’t do, and you can’t switch these off:

  • Invent facts about your game. No made-up numbers, no guessed mechanics. See grounded answers.
  • Give account or purchase support. Those escalate to your team instead.
  • Produce unsafe or abusive content. Standard safety limits apply, regardless of how a player phrases the ask.
  • Leak its own configuration. It won’t repeat its full instructions or expose settings on request.

Handling abuse

Players will test the agent. It’s built to shrug off attempts to bait it, jailbreak it, or make it say something off-brand, and to stay in character while doing so. If someone is spamming it to disrupt the channel, that’s a moderation matter for your team: the agent won’t warn or ban, but you can pause the channel or handle the player with your usual tools.

Rate limits

The agent won’t flood a channel. It paces itself, and it won’t answer the same question over and over for the same player. In a fast channel this keeps it from drowning out people. If you want it to hold back more, or you’re seeing it reply more than you’d like, tell us and we’ll tune the pace for that channel.

Flag controls

You decide how your team flags a bad answer. Choose which emoji reaction counts as a flag, or turn reaction-flagging off, and the report action is always available. See flagging a wrong answer.