Schedule demo

Amazon Bedrock Monitoring


Amazon Bedrock - Overview

Amazon Bedrock is a fully managed service that provides access to high-performing foundation models (FMs) from leading AI companies through a single API. It enables developers to build and scale generative AI applications using techniques such as fine-tuning, retrieval-augmented generation (RAG), and agent orchestration - without managing the underlying infrastructure.

Monitoring Amazon Bedrock is essential for maintaining application performance, controlling costs, and ensuring reliable AI inference operations. Applications Manager's Amazon Bedrock monitoring tool provides real-time visibility into key metrics such as invocation latency, token consumption, throttling events, and log delivery status, enabling you to detect performance bottlenecks, track quota usage, and ensure successful delivery of model invocation logs.

Creating a new Amazon Bedrock monitor

To learn how to create a new Amazon Bedrock monitor, refer here.

Monitored Parameters

Go to the Monitors Category View by clicking the Monitors tab. Click on the Bedrock instance available under Amazon in the Cloud Apps section. Displayed below is the Amazon Bedrock bulk configuration view distributed into three tabs:

  • Availability tab gives the availability history for the past 24 hours or 30 days.
  • Performance tab gives the health status and events for the past 24 hours or 30 days.
  • List view tab enables you to perform bulk admin configurations.

By clicking a monitor from the list, you'll be taken to the Amazon Bedrock dashboard, which includes the following tabs:

Performance Overview

ParameterDescription
SUMMARY
Total AgentsThe total number of agents configured for the Amazon Bedrock monitor.
Failed Agent DeploymentsThe total number of failed Amazon Bedrock agent deployments.
Total Knowledge BasesThe total number of knowledge bases configured for the Amazon Bedrock monitor.
Total Guardrail ProfilesThe total number of guardrail profiles configured for the Amazon Bedrock monitor.
REQUEST TRAFFIC
Requests RateThe total number of Amazon Bedrock model requests per minute between poll intervals, where spikes indicate your infrastructure scales with demand (in requests/min).
Total RequestsThe total number of Amazon Bedrock model requests between poll intervals, where spikes indicate your infrastructure scales with demand.
ERROR REQUESTS BREAKDOWN
Model Client ErrorsThe total number of client-side errors returned by model invocations between poll intervals, indicating invalid user inputs or API call payloads.
Model Server ErrorsThe total number of internal server errors returned by Amazon Bedrock between poll intervals, which helps identify platform-level service degradations.
Throttled RequestsThe total number of throttled model requests between poll intervals, highlighting where request rates have breached account concurrency or rate limits.
TOKENS PER MINUTE (TPM) QUOTA USAGE
TPM Quota UsageThe average amount of Tokens Per Minute (TPM) quota currently consumed between poll intervals, aiding proactive scaling before hitting structural execution thresholds.
REQUEST LATENCY
Request LatencyThe average time taken to process and return model request payloads between poll intervals, where sudden increases signal bottlenecked model operations.
PROMPT TOKENS (INPUT)
Prompt Tokens (Input)The total number of input tokens sent inside prompt bodies between poll intervals, crucial for auditing contextual processing overhead and raw resource costs.
RESPONSE TOKENS (OUTPUT)
Response Tokens (Output)The total number of output tokens generated by model responses between poll intervals, serving as a primary operational metric for billing and throughput.
TIME TO FIRST TOKEN (TTFT)
Time to First Token (TTFT)The average time taken for the model to generate the first response token between poll intervals, mapping directly to perceived end-user latency (in ms).
GENERATED IMAGES
Generated ImagesThe total number of output images generated by multi-modal workflows between poll intervals, helping track downstream image generation pipeline throughput.
PROMPT CACHE PERFORMANCE
Prompt Cache Hit TokensThe total number of input prompt tokens served from cache storage between poll intervals, validating cache hit rates and optimization efficiency.
Prompt Cache Miss TokensThe total number of new input prompt tokens saved into cache memory between poll intervals, representing cold cache misses requiring a backend write loop.

Models

Note: Data collection for Bedrock Models is enabled by default, with a default polling interval of 60 minutes. To modify this, navigate to Settings > Performance Polling. In the Optimize Data Collection tab, select Amazon Bedrock as the Monitor Type, choose Bedrock Models as the Metric Name, and update the Default Polling Status and Polling Interval as required.

ParameterDescription
AI Model Performance Details
Model IDThe model identifier indicating the AI model used by the agent.
Requests RateThe total number of requests routed to this specific model per minute between poll intervals, identifying which foundation units are experiencing the highest consumption (in requests/min).
Total RequestsThe total number of requests routed to this specific model between poll intervals, identifying which foundation units are experiencing the highest consumption.
Model Client ErrorsThe total number of client-side integration faults recorded for this model between poll intervals, pointing to systemic bad prompt requests or mismatched configurations.
Model Server ErrorsThe total number of backend server errors experienced by this specific model between poll intervals, indicating localized service interruptions.
Throttled RequestsThe total number of restricted transactions on this model due to exceeded structural constraints between poll intervals, indicating that the request rate exceeded the model's quota.
Legacy Model RequestsThe total number of deprecated or legacy style invocations processed by this model version between poll intervals, useful for migration tracking.
Request LatencyThe average round-trip latency for requests hitting this specific model identifier between poll intervals, highlighting performance differences across model variants (in ms).
Generated ImagesThe total number of distinct image outputs generated by this specific model between poll intervals, illustrating multi-modal load distributions.
AI Model Token & Quota Details
Model IDThe model identifier indicating the AI model used by the agent.
Prompt Tokens (Input)The total volume of prompt ingestion tokens consumed by this specific model between poll intervals, used to budget individual application usage.
Response Tokens (Output)The total volume of response tokens generated by this specific model context between poll intervals, tracking production billing footprints.
Time To First Token(TTFT)The average time taken for the model to generate the first response token between poll intervals, mapping directly to perceived end-user latency (in ms).
TPM Quota UsageThe average estimated Tokens Per Minute (TPM) allowance consumed by this distinct model configuration between poll intervals, alerting before hitting model ceiling targets.
PERFORMANCE CHARTS
Top 5 Models by Client ErrorsDisplays the top 5 models by the number of client errors between poll intervals.
Top 5 Models by Server ErrorsDisplays the top 5 models by the number of server errors between poll intervals.
Top 5 Models by Request LatencyDisplays the top 5 models by average request latency between poll intervals.
Top 5 Models by Throttling RatesDisplays the top 5 models by the number of throttled requests between poll intervals.

AI Agents

ParameterDescription
Agent Configuration Inventory
Agent IDThe unique technical identifier generated by AWS for this Bedrock Agent instance.
Agent NameThe name configured for the Bedrock agent.
DescriptionThe user description details the agent's configuration.
Latest VersionThe current active agent version identifier of the agent at the time of polling.
Last Updation TimeThe precise timestamp marking the last modification made to this resource or deployment path.
Guardrail IDThe identifier of the active guardrail associated with this operational agent at the time of polling.
Guardrail VersionThe exact structural version of the guardrail profile coupled to this agent lifecycle.
Agent StatusThe operational availability status of the designated Bedrock Agent at the time of polling.

Knowledge Bases

ParameterDescription
KNOWLEDGE BASE DETAILS
Knowledge Base IDThe distinct identifier tracking this vector storage integration.
NameThe human-readable display label configured for this Knowledge Base instance.
DescriptionThe architectural statement documenting this vector resource's indexing purpose.
Last Updation TimeThe timestamp showing the last structural data synchronization or configuration commit on this Knowledge Base.
StatusThe baseline synchronization or provisioning readiness phase of the Knowledge Base at the time of polling.

Guardrails

Note: Data collection for Guardrail Performance is disabled by default. To enable/disable data collection, navigate to Settings > Performance Polling. In the Optimize Data Collection tab, select Amazon Bedrock as the Monitor Type, choose Guardrail Performance as the Metric Name, and update the Default Polling Status as required.

ParameterDescription
Guardrail Details
Guardrail IDThe unique identifier tracking this content moderation and safety profile.
Guardrail NameThe administrative name tracking this content inspection profile.
DescriptionThe user description configured for the guardrail.
VersionThe exact structural version of the guardrail profile coupled to this agent lifecycle.
Creation TimeThe creation timestamp indicating when this functional asset was originally launched inside the region.
Last Updation TimeThe revision timestamp indicating when this profile's rules or limits were last updated.
Cross-Region Routing Profile IDThe regional profile mapping used to route guardrail parameters across regions.
Overall Guardrail Performance Metrics
OperationThe type of transactional API hook currently being inspected by the guardrail policy at the time of polling.
Evaluation Context SourceThe orientation origin of the context processing block being evaluated, tracking whether it is user input or model output at the time of polling.
RequestsThe total number of processing cycles handled through this guardrail intercept loop between the poll interval, helping measure safety checkpoint load profiles.
Client ErrorsThe total number of client interface issues logged against guardrail execution boundaries between poll intervals.
Server ErrorsThe total number of internal system execution crashes during safety scanning operations between the poll interval, pointing to potential policy misconfigurations.
Throttled RequestsThe total number of transactions locked out due to high-volume concurrency limits on guardrail evaluations between poll intervals.
Request Latency(ms)The average processing overhead duration added by executing this safety filter structure between the poll interval, marking the performance footprint of your safety compliance.
Interrupted RequestsThe total number of user interactions blocked, modified, or redacted due to policy compliance hits between the poll interval, displaying the direct mitigation impact of your guardrail setup.
Text UnitsThe total number of text units evaluated by the guardrail filter system between the poll interval, providing a key measure for tracing backend evaluation volumes.
Guardrail Trend Charts
Guardrail Client Errors by Evaluation ContextShows the trend of client interface issues logged against guardrail execution boundaries, broken down by whether the evaluation context was user input or model output.
Guardrail Server Processing Errors by ContextShows the trend of internal system execution crashes during safety scanning operations, broken down by whether the evaluation context was user input or model output.
Guardrail Policy Throttled Evaluations by SourceShows the trend of transactions locked out due to high-volume concurrency limits on guardrail evaluations, broken down by whether the evaluation context was user input or model output.
Average Evaluation Overhead Latency by Content SourceShows the trend of average processing overhead added by executing the safety filter structure, broken down by whether the evaluation context was user input or model output.

Log Delivery Pipelines

ParameterDescription
CLOUDWATCH LOG DESTINATION STATE
Delivered CloudWatch LogsThe total number of request log events successfully written to Amazon CloudWatch Logs between poll intervals, showing stable visibility pipeline execution.
Failed CloudWatch LogsThe total number of request logging failures targeting CloudWatch Destinations between poll intervals, alerting you to identity permissions faults or full target limits.
AMAZON S3 LOG DESTINATION STATE
Delivered S3 LogsThe total number of auditing trail files delivered cleanly into Amazon S3 Buckets between poll intervals, affirming compliance history retention.
Failed S3 LogsThe total number of stream dumps that failed to sync to S3 targets between the poll interval, pointing to configuration issues like broken bucket ACLs or missing write privileges.
Delivered Large Data S3 LogsThe total number of large multi-part payload telemetry files exported successfully to S3 storage targets between poll intervals.
Failed Large Data S3 LogsThe total number of large multi-part payload telemetry transfers that failed to write to Amazon S3 between poll intervals.

Loved by customers all over the world

"Standout Tool With Extensive Monitoring Capabilities"

★★★★★

It allows us to track crucial metrics such as response times, resource utilization, error rates, and transaction performance. The real-time monitoring alerts promptly notify us of any issues or anomalies, enabling us to take immediate action.

Reviewer Role: Research and Development

carlos-rivero
"I like Applications Manager because it helps us to detect issues present in our servers and SQL databases."
Carlos Rivero

Tech Support Manager, Lexmark

Trusted by thousands of leading businesses globally