Amazon Bedrock is a fully managed service that provides access to high-performing foundation models (FMs) from leading AI companies through a single API. It enables developers to build and scale generative AI applications using techniques such as fine-tuning, retrieval-augmented generation (RAG), and agent orchestration - without managing the underlying infrastructure.
Monitoring Amazon Bedrock is essential for maintaining application performance, controlling costs, and ensuring reliable AI inference operations. Applications Manager's Amazon Bedrock monitoring tool provides real-time visibility into key metrics such as invocation latency, token consumption, throttling events, and log delivery status, enabling you to detect performance bottlenecks, track quota usage, and ensure successful delivery of model invocation logs.
To learn how to create a new Amazon Bedrock monitor, refer here.
Go to the Monitors Category View by clicking the Monitors tab. Click on the Bedrock instance available under Amazon in the Cloud Apps section. Displayed below is the Amazon Bedrock bulk configuration view distributed into three tabs:
By clicking a monitor from the list, you'll be taken to the Amazon Bedrock dashboard, which includes the following tabs:
| Parameter | Description |
|---|---|
| SUMMARY | |
| Total Agents | The total number of agents configured for the Amazon Bedrock monitor. |
| Failed Agent Deployments | The total number of failed Amazon Bedrock agent deployments. |
| Total Knowledge Bases | The total number of knowledge bases configured for the Amazon Bedrock monitor. |
| Total Guardrail Profiles | The total number of guardrail profiles configured for the Amazon Bedrock monitor. |
| REQUEST TRAFFIC | |
| Requests Rate | The total number of Amazon Bedrock model requests per minute between poll intervals, where spikes indicate your infrastructure scales with demand (in requests/min). |
| Total Requests | The total number of Amazon Bedrock model requests between poll intervals, where spikes indicate your infrastructure scales with demand. |
| ERROR REQUESTS BREAKDOWN | |
| Model Client Errors | The total number of client-side errors returned by model invocations between poll intervals, indicating invalid user inputs or API call payloads. |
| Model Server Errors | The total number of internal server errors returned by Amazon Bedrock between poll intervals, which helps identify platform-level service degradations. |
| Throttled Requests | The total number of throttled model requests between poll intervals, highlighting where request rates have breached account concurrency or rate limits. |
| TOKENS PER MINUTE (TPM) QUOTA USAGE | |
| TPM Quota Usage | The average amount of Tokens Per Minute (TPM) quota currently consumed between poll intervals, aiding proactive scaling before hitting structural execution thresholds. |
| REQUEST LATENCY | |
| Request Latency | The average time taken to process and return model request payloads between poll intervals, where sudden increases signal bottlenecked model operations. |
| PROMPT TOKENS (INPUT) | |
| Prompt Tokens (Input) | The total number of input tokens sent inside prompt bodies between poll intervals, crucial for auditing contextual processing overhead and raw resource costs. |
| RESPONSE TOKENS (OUTPUT) | |
| Response Tokens (Output) | The total number of output tokens generated by model responses between poll intervals, serving as a primary operational metric for billing and throughput. |
| TIME TO FIRST TOKEN (TTFT) | |
| Time to First Token (TTFT) | The average time taken for the model to generate the first response token between poll intervals, mapping directly to perceived end-user latency (in ms). |
| GENERATED IMAGES | |
| Generated Images | The total number of output images generated by multi-modal workflows between poll intervals, helping track downstream image generation pipeline throughput. |
| PROMPT CACHE PERFORMANCE | |
| Prompt Cache Hit Tokens | The total number of input prompt tokens served from cache storage between poll intervals, validating cache hit rates and optimization efficiency. |
| Prompt Cache Miss Tokens | The total number of new input prompt tokens saved into cache memory between poll intervals, representing cold cache misses requiring a backend write loop. |
Note: Data collection for Bedrock Models is enabled by default, with a default polling interval of 60 minutes. To modify this, navigate to Settings > Performance Polling. In the Optimize Data Collection tab, select Amazon Bedrock as the Monitor Type, choose Bedrock Models as the Metric Name, and update the Default Polling Status and Polling Interval as required.
| Parameter | Description |
|---|---|
| AI Model Performance Details | |
| Model ID | The model identifier indicating the AI model used by the agent. |
| Requests Rate | The total number of requests routed to this specific model per minute between poll intervals, identifying which foundation units are experiencing the highest consumption (in requests/min). |
| Total Requests | The total number of requests routed to this specific model between poll intervals, identifying which foundation units are experiencing the highest consumption. |
| Model Client Errors | The total number of client-side integration faults recorded for this model between poll intervals, pointing to systemic bad prompt requests or mismatched configurations. |
| Model Server Errors | The total number of backend server errors experienced by this specific model between poll intervals, indicating localized service interruptions. |
| Throttled Requests | The total number of restricted transactions on this model due to exceeded structural constraints between poll intervals, indicating that the request rate exceeded the model's quota. |
| Legacy Model Requests | The total number of deprecated or legacy style invocations processed by this model version between poll intervals, useful for migration tracking. |
| Request Latency | The average round-trip latency for requests hitting this specific model identifier between poll intervals, highlighting performance differences across model variants (in ms). |
| Generated Images | The total number of distinct image outputs generated by this specific model between poll intervals, illustrating multi-modal load distributions. |
| AI Model Token & Quota Details | |
| Model ID | The model identifier indicating the AI model used by the agent. |
| Prompt Tokens (Input) | The total volume of prompt ingestion tokens consumed by this specific model between poll intervals, used to budget individual application usage. |
| Response Tokens (Output) | The total volume of response tokens generated by this specific model context between poll intervals, tracking production billing footprints. |
| Time To First Token(TTFT) | The average time taken for the model to generate the first response token between poll intervals, mapping directly to perceived end-user latency (in ms). |
| TPM Quota Usage | The average estimated Tokens Per Minute (TPM) allowance consumed by this distinct model configuration between poll intervals, alerting before hitting model ceiling targets. |
| PERFORMANCE CHARTS | |
| Top 5 Models by Client Errors | Displays the top 5 models by the number of client errors between poll intervals. |
| Top 5 Models by Server Errors | Displays the top 5 models by the number of server errors between poll intervals. |
| Top 5 Models by Request Latency | Displays the top 5 models by average request latency between poll intervals. |
| Top 5 Models by Throttling Rates | Displays the top 5 models by the number of throttled requests between poll intervals. |
| Parameter | Description |
|---|---|
| Agent Configuration Inventory | |
| Agent ID | The unique technical identifier generated by AWS for this Bedrock Agent instance. |
| Agent Name | The name configured for the Bedrock agent. |
| Description | The user description details the agent's configuration. |
| Latest Version | The current active agent version identifier of the agent at the time of polling. |
| Last Updation Time | The precise timestamp marking the last modification made to this resource or deployment path. |
| Guardrail ID | The identifier of the active guardrail associated with this operational agent at the time of polling. |
| Guardrail Version | The exact structural version of the guardrail profile coupled to this agent lifecycle. |
| Agent Status | The operational availability status of the designated Bedrock Agent at the time of polling. |
| Parameter | Description |
|---|---|
| KNOWLEDGE BASE DETAILS | |
| Knowledge Base ID | The distinct identifier tracking this vector storage integration. |
| Name | The human-readable display label configured for this Knowledge Base instance. |
| Description | The architectural statement documenting this vector resource's indexing purpose. |
| Last Updation Time | The timestamp showing the last structural data synchronization or configuration commit on this Knowledge Base. |
| Status | The baseline synchronization or provisioning readiness phase of the Knowledge Base at the time of polling. |
Note: Data collection for Guardrail Performance is disabled by default. To enable/disable data collection, navigate to Settings > Performance Polling. In the Optimize Data Collection tab, select Amazon Bedrock as the Monitor Type, choose Guardrail Performance as the Metric Name, and update the Default Polling Status as required.
| Parameter | Description |
|---|---|
| Guardrail Details | |
| Guardrail ID | The unique identifier tracking this content moderation and safety profile. |
| Guardrail Name | The administrative name tracking this content inspection profile. |
| Description | The user description configured for the guardrail. |
| Version | The exact structural version of the guardrail profile coupled to this agent lifecycle. |
| Creation Time | The creation timestamp indicating when this functional asset was originally launched inside the region. |
| Last Updation Time | The revision timestamp indicating when this profile's rules or limits were last updated. |
| Cross-Region Routing Profile ID | The regional profile mapping used to route guardrail parameters across regions. |
| Overall Guardrail Performance Metrics | |
| Operation | The type of transactional API hook currently being inspected by the guardrail policy at the time of polling. |
| Evaluation Context Source | The orientation origin of the context processing block being evaluated, tracking whether it is user input or model output at the time of polling. |
| Requests | The total number of processing cycles handled through this guardrail intercept loop between the poll interval, helping measure safety checkpoint load profiles. |
| Client Errors | The total number of client interface issues logged against guardrail execution boundaries between poll intervals. |
| Server Errors | The total number of internal system execution crashes during safety scanning operations between the poll interval, pointing to potential policy misconfigurations. |
| Throttled Requests | The total number of transactions locked out due to high-volume concurrency limits on guardrail evaluations between poll intervals. |
| Request Latency(ms) | The average processing overhead duration added by executing this safety filter structure between the poll interval, marking the performance footprint of your safety compliance. |
| Interrupted Requests | The total number of user interactions blocked, modified, or redacted due to policy compliance hits between the poll interval, displaying the direct mitigation impact of your guardrail setup. |
| Text Units | The total number of text units evaluated by the guardrail filter system between the poll interval, providing a key measure for tracing backend evaluation volumes. |
| Guardrail Trend Charts | |
| Guardrail Client Errors by Evaluation Context | Shows the trend of client interface issues logged against guardrail execution boundaries, broken down by whether the evaluation context was user input or model output. |
| Guardrail Server Processing Errors by Context | Shows the trend of internal system execution crashes during safety scanning operations, broken down by whether the evaluation context was user input or model output. |
| Guardrail Policy Throttled Evaluations by Source | Shows the trend of transactions locked out due to high-volume concurrency limits on guardrail evaluations, broken down by whether the evaluation context was user input or model output. |
| Average Evaluation Overhead Latency by Content Source | Shows the trend of average processing overhead added by executing the safety filter structure, broken down by whether the evaluation context was user input or model output. |
| Parameter | Description |
|---|---|
| CLOUDWATCH LOG DESTINATION STATE | |
| Delivered CloudWatch Logs | The total number of request log events successfully written to Amazon CloudWatch Logs between poll intervals, showing stable visibility pipeline execution. |
| Failed CloudWatch Logs | The total number of request logging failures targeting CloudWatch Destinations between poll intervals, alerting you to identity permissions faults or full target limits. |
| AMAZON S3 LOG DESTINATION STATE | |
| Delivered S3 Logs | The total number of auditing trail files delivered cleanly into Amazon S3 Buckets between poll intervals, affirming compliance history retention. |
| Failed S3 Logs | The total number of stream dumps that failed to sync to S3 targets between the poll interval, pointing to configuration issues like broken bucket ACLs or missing write privileges. |
| Delivered Large Data S3 Logs | The total number of large multi-part payload telemetry files exported successfully to S3 storage targets between poll intervals. |
| Failed Large Data S3 Logs | The total number of large multi-part payload telemetry transfers that failed to write to Amazon S3 between poll intervals. |
It allows us to track crucial metrics such as response times, resource utilization, error rates, and transaction performance. The real-time monitoring alerts promptly notify us of any issues or anomalies, enabling us to take immediate action.
Reviewer Role: Research and Development