Property Finder Revolutionizes Incident Management with Autonomous AWS DevOps Agent Integration to Achieve 14-Minute Resolution Times

When a critical production service experiences a sudden CPU spike at 1:00 AM, the traditional incident response lifecycle often descends into a frantic, high-stakes race against the clock. For Property Finder, the leading property portal across the Middle East and North Africa (MENA), such technological disruptions represent more than just technical debt—they signify failed user searches, diminished platform trust, and an immediate negative impact on revenue during high-traffic intervals. Historically, the company, which serves millions of users across five distinct markets, faced a common industry challenge: the time-consuming, manual labor required to correlate metrics and diagnose infrastructure failures. However, by integrating the AWS DevOps Agent into their microservices architecture, Property Finder has successfully transitioned from reactive troubleshooting to an autonomous incident management model, reducing their mean time to resolution (MTTR) from days to a mere 14 minutes.

The Traditional Bottleneck: Manual Incident Response
Before the implementation of the current autonomous pipeline, the incident response workflow at Property Finder was hampered by significant operational latency. When a system alert was triggered, an on-call engineer would be paged, requiring them to manually parse through disparate monitoring tools to correlate metrics—a process that typically consumed 20 to 40 minutes of critical time. During this investigation phase, the platform would often remain in a degraded state, as the root cause remained obscured by the sheer volume of telemetry data.
The business impact of these delays was substantial. Running a distributed microservices architecture on Amazon Elastic Container Service (ECS) and Application Load Balancers (ALB) requires high availability to maintain search performance. When infrastructure issues arose, the resulting service instability prevented real estate agents from updating listings and users from accessing property data, creating a direct correlation between system downtime and financial performance.

A Chronology of Resolution: The 14-Minute Lifecycle
The efficacy of the new autonomous framework was recently validated during a significant production incident involving the company’s core property search microservice. At 1:21 AM, a CloudWatch alarm signaled that the ECS CPU utilization had exceeded 98 percent. Under the legacy model, this would have been the beginning of a multi-hour ordeal. Instead, the incident was resolved entirely within 14 minutes through a highly orchestrated sequence of events.
By 1:22 AM, the investigation had already commenced, and a Slack notification was generated. Simultaneously, four parallel subagents were deployed to analyze logs and metrics, effectively performing the heavy lifting that previously required human intervention. At 1:32 AM, the system successfully identified the root cause: a configuration conflict between CPU and memory target-tracking autoscaling policies. By 1:33 AM, a Jira ticket had been automatically generated, and at 1:34 AM, the on-call engineer was contacted via phone, provided with the complete investigation context. By 1:35 AM, a GitHub pull request (PR) containing the necessary code fix was already open and awaiting review.

Root Cause Analysis: The Danger of Policy Conflicts
The investigation revealed a common yet insidious issue in cloud infrastructure management: conflicting autoscaling policies. The microservice in question was governed by both a CPU target-tracking policy (set at 70 percent) and a memory target-tracking policy (set at 75 percent). While the CPU was under heavy load, the memory usage remained consistently low, at approximately 3 to 8 percent.
This configuration mismatch created a "scaling flip" phenomenon, where the system attempted to reconcile contradictory instructions from the two policies. The instability resulted in 22 separate scaling fluctuations, which effectively prevented the service from adding the necessary capacity to handle the load. The agent identified this logic error in 10 minutes—a task that typically demands a senior engineer with deep expertise in cloud architecture and hours of manual metric correlation.

The Remediation Architecture: Safety by Design
A critical component of Property Finder’s success is the separation of concerns within their remediation layer. The company utilizes a dedicated remediation agent, the pr-creation-agent, which is only invoked after the investigative phase is complete. This architectural choice is central to the safety of the system.
The AWS DevOps Agent operates within a restricted, read-only environment, ensuring that it cannot inadvertently alter infrastructure configurations without oversight. The remediation agent, by contrast, is scoped specifically to interact with the GitHub Management Control Plane (MCP) server. To ensure human-in-the-loop governance, the agent is restricted from auto-merging changes. Every proposed fix is generated as a draft pull request, which engineers must review, validate in a staging environment, and manually merge. This ensures that while the speed of diagnosis and preparation is automated, the final decision-making authority remains with the human engineering team.

Strategic Implications and Quantitative Gains
The shift to autonomous incident management has fundamentally altered the operational landscape at Property Finder. The quantitative improvements are striking: an 88 percent reduction in end-to-end incident resolution time, a 50 to 75 percent decrease in investigation time, and the elimination of manual documentation burdens. Furthermore, the automation of root cause analysis ensures that every incident is comprehensively logged, providing a historical record that was previously incomplete or manually documented.
Yasitha Bogamuwa, Cloud Engineering Manager at Property Finder, highlighted the strategic value of this transition: "We now rely fully on the AWS DevOps Agent to identify infrastructure-related issues. It has helped us identify multiple complex issues without even opening a support ticket. Even if we had raised tickets, it would likely have taken support engineers hours to find the root cause, whereas we resolved these issues in minutes."

Broader Industry Impact
The implications for the broader tech industry are significant. As distributed architectures become increasingly complex, the human capacity to monitor and troubleshoot in real-time is being stretched to its limits. Property Finder’s implementation serves as a blueprint for organizations looking to leverage generative AI and autonomous agents to maintain service reliability at scale.
By integrating the AWS DevOps Agent with existing tools like Jira, Grafana IRM, and GitHub, the company has created an ecosystem where the burden of "alert fatigue" is shifted from the engineer to the machine. The result is not just a faster resolution time, but a higher quality of life for on-call engineers, who are no longer forced to perform complex forensic analysis in the middle of the night.

Moving Forward: The Future of Autonomous DevOps
While the current implementation at Property Finder is highly successful, it is indicative of a broader trend toward autonomous systems in software engineering. The reliance on EventBridge for orchestrating these workflows—triggering downstream actions only when investigation is confirmed—demonstrates a mature approach to event-driven architecture.
For other organizations seeking to replicate this success, the prerequisites involve configuring robust webhook triggers for alarms and establishing clear, event-driven outputs. As the industry continues to refine these autonomous models, the role of the DevOps engineer is likely to evolve from that of a manual "firefighter" to an architect of automation, designing the guardrails and governance structures that allow these agents to operate safely and effectively. The Property Finder case study confirms that when designed with safety and human oversight at its core, autonomous incident management is not just a future goal—it is a present-day reality that is redefining the standards of operational excellence in the cloud.







