OneOps Enables Secure, Resilient Enterprise AI Governance and Operational Intelligence on AWS with OneData
Learn how OneOps partnered with OneData Software Solutions to build a centralized AI governance platform on AWS, enabling secure LLM adoption through intelligent governance, contextual AI, cost optimization, scalable infrastructure, operational monitoring, data protection, and AWS CDK-based disaster recovery readiness.
Benefits
Centralized governance for enterprise-wide LLM usage
Improved visibility into prompts, responses, user activity, model utilization, and token consumption
Enhanced security through prompt validation, access controls, and enterprise data protection guardrails
Optimized AI costs through token-level monitoring and departmental cost attribution
Delivered context-aware AI responses using Retrieval-Augmented Generation
Enabled automated governance and operational intelligence through Agentic AI
Improved application scalability using Kubernetes Horizontal Pod Autoscaler, EKS managed node groups, and Amazon EC2 Auto Scaling
Strengthened operational visibility through Amazon CloudWatch monitoring
Improved data durability through Amazon S3-based storage
Protected PostgreSQL data through encrypted daily backups with seven-day retention
Established a four-hour Recovery Time Objective and a 24-hour Recovery Point Objective
Enabled repeatable infrastructure reconstruction using AWS CDK-based Infrastructure as Code
Enabled consistent application redeployment using Amazon ECR and GitLab CI/CD
Established an estimated recovery and redeployment duration of approximately three to four hours
Created a technical foundation for future formal disaster recovery testing
About the Customer
OneOps is an enterprise AI governance and operational intelligence platform designed to help organizations securely adopt, govern, and manage Large Language Models at scale. The platform enables enterprises to centralize AI governance, control model access, monitor prompts and user activity, optimize AI costs, protect sensitive organizational information, and improve visibility across enterprise AI workloads.
Overview
As enterprise adoption of Generative AI accelerated, OneOps identified the need for a centralized governance platform capable of securely managing Large Language Models across multiple business teams, departments, and AI providers.
Without a unified governance layer, AI usage became fragmented across the organization. Different teams independently adopted and interacted with AI models, making it difficult to maintain visibility, enforce access policies, control token consumption, attribute costs, protect sensitive enterprise information, and support internal governance requirements.
The organization also required a resilient technical foundation capable of scaling with increasing usage, protecting enterprise knowledge, monitoring application health, preserving audit records, maintaining database backups, and restoring the platform following a major infrastructure disruption.
To address these requirements, OneOps partnered with OneData Software Solutions to design and develop OneOps.pro, an AWS-powered AI governance and operational intelligence platform.
The solution combines Amazon Bedrock, Amazon Elastic Kubernetes Service, Amazon Elastic Container Registry, Amazon S3, Amazon CloudWatch, Elastic Load Balancing, AWS Certificate Manager, AWS security and audit services, PostgreSQL, GitLab CI/CD, and AWS CDK-based Infrastructure as Code.
The implementation provides centralized AI governance, contextual intelligence, operational monitoring, scalable containerized infrastructure, security controls, durable storage, data protection, and repeatable recovery within a unified enterprise AI environment.
Overview
As enterprise adoption of Generative AI accelerated, OneOps identified the need for a centralized governance platform capable of securely managing Large Language Models across multiple business teams, departments, and AI providers.
Without a unified governance layer, AI usage became fragmented across the organization. Different teams independently adopted and interacted with AI models, making it difficult to maintain visibility, enforce access policies, control token consumption, attribute costs, protect sensitive enterprise information, and support internal governance requirements.
The organization also required a resilient technical foundation capable of scaling with increasing usage, protecting enterprise knowledge, monitoring application health, preserving audit records, maintaining database backups, and restoring the platform following a major infrastructure disruption.
To address these requirements, OneOps partnered with OneData Software Solutions to design and develop OneOps.pro, an AWS-powered AI governance and operational intelligence platform.
The solution combines Amazon Bedrock, Amazon Elastic Kubernetes Service, Amazon Elastic Container Registry, Amazon S3, Amazon CloudWatch, Elastic Load Balancing, AWS Certificate Manager, AWS security and audit services, PostgreSQL, GitLab CI/CD, and AWS CDK-based Infrastructure as Code.
The implementation provides centralized AI governance, contextual intelligence, operational monitoring, scalable containerized infrastructure, security controls, durable storage, data protection, and repeatable recovery within a unified enterprise AI environment.
Opportunity | Establishing Secure and Resilient Enterprise AI Governance
As organizations rapidly adopted Generative AI technologies, OneOps recognized the need to establish centralized governance and operational intelligence across enterprise AI usage.
Multiple business teams independently adopted different Large Language Models without centralized oversight, resulting in fragmented AI usage, inconsistent governance policies, limited operational visibility, and reduced accountability.
The organization lacked structured access controls capable of defining which users, departments, or business units could access specific models. Prompts, responses, user activity, and model interactions were not centrally governed, making auditability and policy enforcement increasingly difficult.
Operational costs also became challenging to manage. Token-based AI consumption was not centrally monitored, resulting in unpredictable spending and limited ability to attribute costs to individual users, departments, or business units.
Users also required AI responses grounded in internal organizational knowledge rather than generic public information. Without contextual intelligence, AI-generated responses could lack business relevance, accuracy, and consistency.
Security and compliance presented additional challenges. Sensitive enterprise information could potentially be exposed through AI interactions without prompt validation, role-based access controls, and data protection guardrails. The lack of centralized monitoring and audit trails also complicated security investigations, compliance reviews, and governance reporting.
The platform needed to support changing workload demand. As the number of users, prompts, model interactions, and governance events increased, the application environment needed to scale without extensive manual intervention.
From a resilience perspective, OneOps required a platform capable of:
- Scaling application and compute capacity according to workload demand
- Protecting application, governance, and enterprise knowledge data
- Maintaining encrypted PostgreSQL backups
- Preserving audit and operational records
- Monitoring application and infrastructure health
- Supporting repeatable application deployment
- Recreating the AWS environment following a major infrastructure failure
- Restoring the platform within a four-hour Recovery Time Objective
- Limiting potential data loss to a 24-hour Recovery Point Objective
The current platform operates within a single Availability Zone. Therefore, the resilience strategy focuses on rapid infrastructure reconstruction, containerized application recovery, encrypted data backups, monitoring, and operational restoration rather than automatic Multi-AZ failover.
A manual infrastructure reconstruction process could increase recovery time, introduce configuration inconsistencies, create deployment errors, and delay restoration of critical AI governance services.
These challenges highlighted the need for a centralized platform capable of combining AI governance, contextual intelligence, cost control, security, scalable infrastructure, operational monitoring, data protection, and disaster recovery readiness.
Solution | Implementing an AWS-Powered Enterprise AI Governance Platform
To address these challenges, OneData designed and developed OneOps.pro, a centralized AI governance and operational intelligence platform built on AWS.
The implementation began with an assessment of the organization’s AI usage, governance requirements, access policies, cost visibility challenges, security risks, application dependencies, and recovery requirements.
Based on these findings, OneData designed a containerized AWS environment capable of controlling AI access, monitoring usage, protecting enterprise information, scaling workloads, and supporting infrastructure and application recovery.
The platform’s Generative AI capabilities are powered by Amazon Bedrock. This enables natural language interaction, intelligent summaries, contextual insights, and flexible access to supported foundation models according to organizational requirements.
Retrieval-Augmented Generation was implemented to integrate approved internal enterprise knowledge into AI interactions. Relevant documents, knowledge repositories, application assets, and organizational information are retrieved and supplied to the model, enabling more accurate and context-aware responses.
Amazon S3 provides durable object storage for enterprise knowledge, platform assets, and information used by the RAG workflow. Separating durable data from individual application containers allows relevant information to remain available if the application environment must be recreated.
A FastAPI-based integration layer coordinates interactions between the user interface, governance controls, token management, enterprise data, and Amazon Bedrock. It validates access, applies policies, enforces token limitations, routes requests to approved models, and returns governed responses to users.
A centralized governance layer was implemented to control access to models and platform functions according to user role, department, organization, approved use case, and governance policy.
Organization- and department-level token controls enable enterprises to define consumption limits, improve cost attribution, prevent uncontrolled usage, and identify unusual token-consumption patterns.
The platform captures relevant AI activity, including prompts, responses, user activity, model invocations, token consumption, governance events, and policy decisions. These records improve transparency, accountability, cost analysis, governance reporting, and auditability.
An Agentic AI layer supports automated governance and operational intelligence. Intelligent agents monitor usage patterns, identify anomalies, detect excessive consumption, generate alerts, enforce approved policies, and support corrective actions.
The application and PostgreSQL database are hosted within Amazon Elastic Kubernetes Service. Amazon EKS provides a consistent containerized deployment environment for managing application and database workloads.
Kubernetes Horizontal Pod Autoscaler adjusts application pod capacity according to workload demand. EKS managed node groups and Amazon EC2 Auto Scaling provide scalable compute capacity for the Kubernetes environment.
This enables the platform to respond to changes in user activity, prompt volume, model requests, token consumption, and governance workloads without relying entirely on manual compute provisioning.
Incoming traffic is managed through Elastic Load Balancing, while AWS Certificate Manager supports encrypted application communications. Application and PostgreSQL workloads operate within private subnet segments, reducing direct public exposure.
Amazon Elastic Container Registry is used to securely store approved application container images. Centralized image management supports version control, traceability, deployment consistency, rollback readiness, and application recovery.
GitLab CI/CD supports the application build, validation, packaging, and deployment process. The pipeline integrates with Amazon ECR and Amazon EKS, enabling approved application versions to be deployed consistently during normal operations and recovery scenarios.
Amazon CloudWatch provides centralized visibility into application and infrastructure operations. Metrics, logs, and operational events support health monitoring, EKS visibility, infrastructure utilization tracking, troubleshooting, alerting, and incident investigation.
Security and operational governance are further strengthened through AWS-native services. AWS Identity and Access Management controls access to AWS resources, Amazon GuardDuty supports continuous threat detection, AWS Security Hub centralizes security findings, Amazon Inspector supports vulnerability management, AWS Config records configuration changes, and AWS CloudTrail captures AWS account activity and API events.
To protect critical application and governance data, OneData implemented a PostgreSQL backup strategy with the following configuration:
- Backups performed every 24 hours
- Backups retained for seven days
- Backup encryption enabled
- Recovery Point Objective established at 24 hours
- Point-in-time recovery is not enabled
The scheduled backup process provides recovery points based on the 24-hour backup cycle. The latest available encrypted backup can be restored following data corruption, accidental deletion, application failure, or infrastructure disruption.
Because point-in-time recovery is not enabled, database restoration is limited to the latest available scheduled backup. This approach supports the defined 24-hour Recovery Point Objective.
To strengthen disaster recovery readiness, OneData maintains the AWS environment using AWS Cloud Development Kit-based Infrastructure as Code.
AWS CDK defines the required network resources, security configurations, Amazon EKS environment, managed node groups, scaling resources, load-balancing components, monitoring services, storage resources, and supporting application dependencies in a repeatable and version-controlled format.
If a major infrastructure failure or disaster affects the environment, AWS CDK can be used to recreate the required AWS services and application infrastructure. This reduces dependence on manual reconstruction, minimizes configuration inconsistencies, and improves deployment standardization.
The documented recovery approach includes recreating the AWS infrastructure using AWS CDK, re-establishing Amazon EKS and its managed node groups, retrieving approved container images from Amazon ECR, redeploying the application through GitLab CI/CD, restoring PostgreSQL from the latest encrypted backup, reconnecting Amazon S3-based enterprise assets, restoring Amazon Bedrock integrations, and validating governance controls, token policies, monitoring, and audit logging.
OneData and OneOps established a Recovery Time Objective of four hours and a Recovery Point Objective of 24 hours.
Based on the documented infrastructure reconstruction, application deployment, and database restoration workflow, the estimated recovery and redeployment duration is approximately three to four hours.
However, a formal end-to-end recovery test has not yet been completed. Therefore, the estimated three-to-four-hour recovery duration and the four-hour RTO have not yet been validated through a documented disaster recovery exercise.
Formal recovery testing has been identified as the next resilience improvement activity.
Outcome | Enabling Secure, Resilient, and Scalable Enterprise AI Adoption
Following the implementation of OneOps.pro, the organization established a centralized AI governance platform capable of securely managing enterprise-wide Large Language Model adoption while improving operational visibility, scalability, security, data protection, financial accountability, and recovery readiness.
Established centralized governance across enterprise-wide AI usage
Improved visibility into prompts, responses, users, model activity, and token consumption
Implemented organization-, department-, and role-based model access controls
Enabled organizational and departmental token limitations
Improved AI cost attribution and usage accountability
Reduced the risk of uncontrolled AI consumption
Delivered context-aware responses using Retrieval-Augmented Generation
Enabled governed access to foundation models using Amazon Bedrock
Enabled proactive governance through Agentic AI
Improved workload scalability through Kubernetes Horizontal Pod Autoscaler
Increased compute scalability through EKS managed node groups and Amazon EC2 Auto Scaling
Enabled consistent application image management through Amazon ECR
Enabled repeatable application deployment and restoration through GitLab CI/CD
Improved data durability through Amazon S3-based storage
Improved operational visibility through Amazon CloudWatch
Strengthened threat detection, vulnerability visibility, configuration monitoring, and auditability through AWS-native security services
Established encrypted PostgreSQL backups every 24 hours
Established a seven-day PostgreSQL backup retention period
Implemented scheduled backup-based recovery without point-in-time recovery
Defined a four-hour Recovery Time Objective
Defined a 24-hour Recovery Point Objective
Established AWS CDK-based Infrastructure as Code for repeatable infrastructure deployment
Enabled consistent redeployment of the application and required AWS services following a major disruption
Reduced dependence on manual infrastructure reconstruction
Minimized recovery-related configuration inconsistencies
Documented an estimated recovery and redeployment duration of approximately three to four hours
Established the technical foundation required for formal recovery testing
Identified end-to-end disaster recovery validation as the next resilience improvement activity
Established a scalable, secure, and recoverable foundation for responsible enterprise AI adoption
By combining Amazon Bedrock, Amazon EKS, scalable container infrastructure, Amazon ECR, Amazon S3, Amazon CloudWatch, PostgreSQL backups, GitLab CI/CD, and AWS CDK-based Infrastructure as Code, OneOps.pro provides a secure and recoverable foundation for enterprise AI governance. The solution is designed around a four-hour RTO and a 24-hour RPO, supported by encrypted daily PostgreSQL backups retained for seven days. In the event of a major disruption, AWS CDK, Amazon ECR, GitLab CI/CD, Amazon S3, and the latest available database backup can be used to rebuild and restore the platform. The estimated recovery duration is three to four hours, although this has not yet been formally validated through an end-to-end recovery test.
Build a more secure and cost-efficient
AWS environment
Partner with OneData to optimize your cloud infrastructure, reduce costs, and
strengthen security—without compromising performance.