/EvaluateToolRegistry
Build an evaluation harness for a registry of tools an agent can call — e.g. a set of ten internal APIs an agent may invoke.
Your AI Command Vault
Searching for the meaning of life… try /RootCause.
Agent Architecture
60 commands
Build an evaluation harness for a registry of tools an agent can call — e.g. a set of ten internal APIs an agent may invoke.
Design the retry and fallback logic for a data-entry automation agent — e.g. an agent that copies invoice fields into the accounting system.
Set up monitoring and alerting for an autonomous research agent — e.g. an agent that compiles competitor pricing every week.
Plan the multi-agent orchestration for a human-in-the-loop hand-off protocol — e.g. escalating a support agent's uncertain answers to a human.
Design the architecture for a scheduling assistant agent — e.g. an agent that negotiates meeting times across three time zones.
Build an evaluation harness for a coding assistant agent — e.g. an agent that opens pull requests to fix flaky tests.
Set up monitoring and alerting for an agent's token and API cost budget — e.g. a daily token cap for an agent running on GPT-4.
Define the safe execution sandbox for an agent's long-term memory store — e.g. conversation history that should persist across sessions.
Plan the multi-agent orchestration for a customer-support agent — e.g. an agent that resolves refund requests under $50 automatically.
Build an evaluation harness for a plan-execute-reflect loop — e.g. an agent that plans steps.
Design the retry and fallback logic for a registry of tools an agent can call — e.g. a set of ten internal APIs an agent may invoke.
Set up monitoring and alerting for a data-entry automation agent — e.g. an agent that copies invoice fields into the accounting system.
Define the safe execution sandbox for an autonomous research agent — e.g. an agent that compiles competitor pricing every week.
Design the architecture for a human-in-the-loop hand-off protocol — e.g. escalating a support agent's uncertain answers to a human.
Build an evaluation harness for a scheduling assistant agent — e.g. an agent that negotiates meeting times across three time zones.
Design the retry and fallback logic for a coding assistant agent — e.g. an agent that opens pull requests to fix flaky tests.
Define the safe execution sandbox for an agent's token and API cost budget — e.g. a daily token cap for an agent running on GPT-4.
Plan the multi-agent orchestration for an agent's long-term memory store — e.g. conversation history that should persist across sessions.
Design the architecture for a customer-support agent — e.g. an agent that resolves refund requests under $50 automatically.
Design the retry and fallback logic for a plan-execute-reflect loop — e.g. an agent that plans steps.
Set up monitoring and alerting for a registry of tools an agent can call — e.g. a set of ten internal APIs an agent may invoke.
Define the safe execution sandbox for a data-entry automation agent — e.g. an agent that copies invoice fields into the accounting system.
Plan the multi-agent orchestration for an autonomous research agent — e.g. an agent that compiles competitor pricing every week.
Build an evaluation harness for a human-in-the-loop hand-off protocol — e.g. escalating a support agent's uncertain answers to a human.
Design the retry and fallback logic for a scheduling assistant agent — e.g. an agent that negotiates meeting times across three time zones.
Set up monitoring and alerting for a coding assistant agent — e.g. an agent that opens pull requests to fix flaky tests.
Plan the multi-agent orchestration for an agent's token and API cost budget — e.g. a daily token cap for an agent running on GPT-4.
Design the architecture for an agent's long-term memory store — e.g. conversation history that should persist across sessions.
Build an evaluation harness for a customer-support agent — e.g. an agent that resolves refund requests under $50 automatically.
Set up monitoring and alerting for a plan-execute-reflect loop — e.g. an agent that plans steps.
Define the safe execution sandbox for a registry of tools an agent can call — e.g. a set of ten internal APIs an agent may invoke.
Plan the multi-agent orchestration for a data-entry automation agent — e.g. an agent that copies invoice fields into the accounting system.
Design the architecture for an autonomous research agent — e.g. an agent that compiles competitor pricing every week.
Design the retry and fallback logic for a human-in-the-loop hand-off protocol — e.g. escalating a support agent's uncertain answers to a human.
Set up monitoring and alerting for a scheduling assistant agent — e.g. an agent that negotiates meeting times across three time zones.
Define the safe execution sandbox for a coding assistant agent — e.g. an agent that opens pull requests to fix flaky tests.
Design the architecture for an agent's token and API cost budget — e.g. a daily token cap for an agent running on GPT-4.
Build an evaluation harness for an agent's long-term memory store — e.g. conversation history that should persist across sessions.
Design the retry and fallback logic for a customer-support agent — e.g. an agent that resolves refund requests under $50 automatically.
Define the safe execution sandbox for a plan-execute-reflect loop — e.g. an agent that plans steps.
Plan the multi-agent orchestration for a registry of tools an agent can call — e.g. a set of ten internal APIs an agent may invoke.
Design the architecture for a data-entry automation agent — e.g. an agent that copies invoice fields into the accounting system.
Build an evaluation harness for an autonomous research agent — e.g. an agent that compiles competitor pricing every week.
Set up monitoring and alerting for a human-in-the-loop hand-off protocol — e.g. escalating a support agent's uncertain answers to a human.
Define the safe execution sandbox for a scheduling assistant agent — e.g. an agent that negotiates meeting times across three time zones.
Plan the multi-agent orchestration for a coding assistant agent — e.g. an agent that opens pull requests to fix flaky tests.
Build an evaluation harness for an agent's token and API cost budget — e.g. a daily token cap for an agent running on GPT-4.
Design the retry and fallback logic for an agent's long-term memory store — e.g. conversation history that should persist across sessions.
Set up monitoring and alerting for a customer-support agent — e.g. an agent that resolves refund requests under $50 automatically.
Plan the multi-agent orchestration for a plan-execute-reflect loop — e.g. an agent that plans steps.
Design the architecture for a registry of tools an agent can call — e.g. a set of ten internal APIs an agent may invoke.
Build an evaluation harness for a data-entry automation agent — e.g. an agent that copies invoice fields into the accounting system.
Design the retry and fallback logic for an autonomous research agent — e.g. an agent that compiles competitor pricing every week.
Define the safe execution sandbox for a human-in-the-loop hand-off protocol — e.g. escalating a support agent's uncertain answers to a human.
Plan the multi-agent orchestration for a scheduling assistant agent — e.g. an agent that negotiates meeting times across three time zones.
Design the architecture for a coding assistant agent — e.g. an agent that opens pull requests to fix flaky tests.
Design the retry and fallback logic for an agent's token and API cost budget — e.g. a daily token cap for an agent running on GPT-4.
Set up monitoring and alerting for an agent's long-term memory store — e.g. conversation history that should persist across sessions.
Define the safe execution sandbox for a customer-support agent — e.g. an agent that resolves refund requests under $50 automatically.
Design the architecture for a plan-execute-reflect loop — e.g. an agent that plans steps.