Monitor your Amazon Bedrock workloads with Applications Manager
Organizations are increasingly integrating GenAI capabilities into their applications to deliver richer, more contextual user experiences—from AI-powered customer support and enterprise search to content generation, virtual assistants, and automated workflows.
To build and scale these GenAI-powered experiences, they are turning to platforms such as Amazon Bedrock, which provides access to foundation models that developers can integrate into their applications.
As these AI-powered experiences become business-critical, they also introduce a new responsibility for IT and cloud teams: keeping the GenAI workloads powering them available, responsive, and within service limits.
An application might be running perfectly well, but what happens when the AI model it depends on starts responding slowly or fails to respond? Any degradation in the AI layer can ultimately affect the experience delivered to the end user.
This makes visibility into the performance and usage of GenAI workloads increasingly important. ManageEngine Applications Manager now supports Amazon Bedrock monitoring, enabling IT and cloud teams to track performance, usage, errors, and quota consumption of their Bedrock workloads alongside the rest of their application environment.
What is Amazon Bedrock?
Amazon Bedrock is a fully managed AWS service that gives organizations access to foundation models from leading AI providers through a unified API. These models can be used to add GenAI capabilities such as conversational assistants, content generation, enterprise search, and intelligent automation to applications.
For organizations already operating within the AWS ecosystem, Bedrock provides a managed way to build and scale these GenAI experiences without having to manage the underlying model infrastructure.
But Bedrock goes beyond simply connecting an application to an AI model. Organizations can customize foundation models for specific needs, use RAG to ground AI responses in their own organizational data, and build AI agents capable of orchestrating multi-step tasks and actions.
For example, a bank could use Bedrock to power an AI assistant that answers customer questions using the bank's latest policies and knowledge resources. An AI agent could go a step further by orchestrating actions across the applications and APIs involved in fulfilling a customer request.
As these capabilities become part of business-critical applications, monitoring just the application is no longer enough. Teams also need visibility into the Bedrock services and AI workloads powering these experiences.
GenAI introduces a new layer to application monitoring
IT teams are accustomed to monitoring application dependencies such as servers, databases, APIs, containers, and cloud services. With GenAI-powered applications, the dependency chain expands.
IT teams now need answers to a different set of questions:
How many model requests are being processed?
Are AI responses getting slower?
How many tokens are being consumed?
Are requests failing?
Are workloads approaching service quotas?
Are requests being throttled?
This is where Amazon Bedrock monitoring in Applications Manager comes in. It helps bring these signals into the broader application monitoring environment.
Let's look at what teams can monitor and, more importantly, what the metrics will tell them.
Supported metrics | What it helps IT teams understand |
Invocations | How many successful requests are being processed by models and whether AI workload usage is increasing or decreasing. |
Invocation latency Time to First Token (TTFT) | How quickly models start responding and how long requests take to complete, helping teams spot slow AI experiences. |
Input tokens Output tokens | How much content models are processing and generating, helping teams track changes in GenAI usage. |
Estimated Tokens Per Minute (TPM) Quota Usage | How close workloads are to their token throughput limits so teams can identify potential capacity issues early. |
Client errors/ Server errors/ Throttled invocations | Whether requests are failing due to request-related issues, service-side problems, or throttling caused by applicable limits. |
Go beyond model monitoring with Amazon Bedrock Agents
Amazon Bedrock provides more than access to foundation models. Organizations can also build AI agents that go beyond generating responses to perform more complex, multi-step tasks.
For example, consider a banking assistant handling this request:
“I lost my credit card. Help me block it and tell me how to get a replacement.”
Responding to this request will involve more than generating an answer. The AI agent needs to understand the customer's intent, invoke an application or API to initiate an action, retrieve the bank's latest card replacement procedure, and then generate an appropriate response.
Amazon Bedrock Agents can orchestrate foundation models with other components, such as APIs and knowledge bases, to carry out multi-step workflows.
Applications Manager extends its Amazon Bedrock visibility to these agent-based workflows, helping teams monitor:
Agent configuration and status: Teams gain visibility into agent details, configurations, and versions.
Knowledge base associations: They track the knowledge bases connected to agents and their association status. Knowledge bases enable agents to retrieve relevant organizational information to provide more contextual responses.
Guardrail associations: They view the guardrails associated with agents and their status. Guardrails help control the inputs and responses of GenAI applications based on defined safeguards.
Agent performance: They monitor requests, latency, errors, throttling, and token consumption to understand how agents are performing.
Model performance: They track the performance of the underlying model invocations powering agent interactions.
By monitoring both Amazon Bedrock and Amazon Bedrock Agents, IT and cloud teams can gain visibility beyond the underlying model calls and understand the health of the broader components powering their GenAI applications.
Key Amazon Bedrock monitoring terms
For teams coming from traditional application or infrastructure monitoring, some GenAI metrics may be unfamiliar. Here are a few key terms:
Invocation: A request made to a foundation model.
Token: A unit of text processed or generated by a foundation model.
Input tokens: Tokens sent to the model as part of the request.
Output tokens: Tokens generated by the model in its response.
Time to First Token (TTFT): How long it takes for the model to begin responding.
Tokens Per Minute (TPM): The number of tokens processed within a minute.
Throttling: Requests being restricted when applicable service limits are reached.
Understanding these signals makes it easier for traditional IT teams to extend their monitoring practices to GenAI workloads.
Bring GenAI monitoring into broader application visibility
An AI-powered business service may depend on application services, APIs, databases, containers, cloud resources, and foundation models, all of which contribute to the final user experience.
Applications Manager's Amazon Bedrock monitoring enables teams to bring visibility into Bedrock workloads alongside the other components supporting their applications.
Rather than monitoring Bedrock separately, organizations can view AI workloads alongside the rest of their AWS infrastructure and application stack from a single platform.
For teams running production AI applications, this means monitoring Bedrock health and performance from the same console used for the rest of the AWS environment. The result is faster troubleshooting, easier correlation across infrastructure and AI services, and more efficient root-cause analysis.
Explore Amazon Bedrock monitoring with ManageEngine Applications Manager.