ArchVault_Multi-Region Resilient Platform Architecture on AWS

Give your colleagues a place to learn about your team and what you’re working on. Use the + buttons in your left-hand sidebar to add more pages, like process docs or a project roadmap.

About the Team

@Isiaka Ismail @abdulhakeemakinsile @Edidiong Ekong @Marvin Unumadu


Phase 1:

Requirements Analysis, Architecture Design, ADRs, etc.

A. Business Requirements Interpretation

Requirement Target Metric Operational Detail AWS Mechanisms Validation Requirement Met
99.95% uptime Max 4.38 hrs per yearequalling 26 min per month Ensure that no single component or service is a point of failure. In absolute terms, this also includes planned maintenance. Multi-Availability Zones across 2 AZs.

Application Load Balancers (ALBs) with health checks.

Elastic Container Service (ECS).

Aurora DBs.

All combined across tiers to deliver automated failover without human interference. | Chaos Engineering: Terminate all AZ-a ECS tasks and ensure error rate stays below 0.5%.

Task count must be restored within 3 minutes.

Uptime will be measured across load tests and failover drill. | Yes ✅ | | Data Residency in Nigeria | All data must be 100% resident within the region (af-south-1) corresponding to Africa, Cape Town. | A mandatory requirement to guarantee the storage and processing of customer data (especially PII) falls within the assigned region.

DR replication must be encrypted-only (”Zero-Knowledge”) | Primary: All services must be created within af-south-1 region.

DR: Aurora Global sends encryption storage pages only to DR region. KMS master stays wilthin af-south-1.

Create Config rules to block resource creation outside the specified region.  | AWS Config Compliance Dashboard for continuous compliance reporting.

Amazon Macie for automated PII scanning on S3 buckets.

VPC Flow Logs + GuardDuty for network egress monitoring to audit external IP endpoints and ensure no unencrypted data flows cross-region. | Conditional Yes ☑️ | | Recovery Time Objective (RTO)  | 15 mins. | Recovery must be automated and completed within 15 mins max. No human intervention (self-healing).

DR region must be pre-warmed (pilot light).

Traffic routing must detect failure within 30s.

DNS TTL must be pre-staged to 60s | AWS Route 53 failover routhing with health check intervals of 10s and a threshold of 2 failures.

TTL of 60s pre-staged up to 24hrs before events.

Aurora DBs promote global availability and limited replication lag.

The ECS Pilot Light DR model helps keep ECS clusters pre-warmed. | A timed failover drill with every step timestamped.

Target traffic routing to DR must fall within 15 - 17 mins max.

| Conditional Yes ☑️ | | Recovery Point Objective (RPO) | 5 mins max. | This is a permissible limit of transaction data loss in the event of a disaster.

Database replication process must be continuous and not batched.

Replication lag must be tracked in seconds and replication mechanism must survive regional-level failure preferably without data loss at moment of failure. | Aurora Global DB with replication lag under 1 second.

Aurora will maintain 6 storage copies across both AZs internally.

Consider enabling point-in-time recovery with a 35-day backup retention. | During the Failover Drill: Commit a test transaction 1 minute before the failure.

After Promotion (DR): Execute a query of DR Aurora for the same transaction ID.

Actual RPO should be the gap between the last committed transaction and the first readable record in DR. | Yes ✅ | | 10x Peak Load | Accommodate 10x of current transactions during end-of-month spikes without any degradation. | Autoscaling should ideally be triggered within 3 minutes of the traffic spike without manual intervention.

99% of response time (P99 Latency) must remain below 2 seconds with the 10x load.

Error rates should also stay below 0.1%. | ECS Fargate with Target-Tracking Auto-Scaling combinining CPU (60%) + ALB request counts will trigger scaling.

Scaleout Cooldown: within 30 seconds.

Max Tasks: 50

ElastiCache Redis caches 80% of the read-heavy paths, while RDS Proxy prevents the Aurora connections from being overwhelmed with a barge of requests.  | A Grafana K6 load test ramped to 10x over 10 minutes and sustained for 20 minutes.

A record of the p50/ p95/ p99 latency, the error rate, ECS task count, and Aurora connection count.

These records must meet set out targets | Yes ✅ | | Data Encryption | All data is encrypted at-rest & in-transit and must utilize customer-managed keys. | Four (4) separate CMKs must be created and rotated annually.

• Aurora • S3 (Invoices) • S3 (Logs) • Secrets Manager

All API calls must be run over TLS 1.2+

S3 bucket policy to deny HTTP must be in place. | KMS Customer Managed Keys (CMKs) with automatic rotation.

Aurora Server Side Encryption (SSE)-KMS.

S3 Server Side Encryption (SSE)-KMS with bucket policy DenyHTTP.

Secrets Manager encrypted with CMK.

ElastiCache in-transit TLS and at-rest KMS

ALB TLS policy.

| AWS Config rule check: Encrypted volumes and S3 bucket SSE enabled.

CloudTrail: To track every KMS decrypt call.

To go an extra mile, we can try an unencrypted PUT request in the S3 bucket which should ideally return a 403 error code. | Yes ✅ | | Clear and Immutable Audit Trail | Every API call, data access, and config adjustment must be permanently logged and queryable. | CloudTrail must be enabled in all regions (multi-region trail).

Log file validation must be enabled

Logs delivered to the S3 buckets must have MFA delete enabled as well as Object Lock.

KMS usage must be logged.

WAF request logs must also be captured. | CloudTrail with multi-region trail and logs submitted to an S3 bucket.

Log File Validation: This can be done utilizing a SHA-256 digest chain.

S3 Object Lock compliance with a 7 year retention policy.

Enable CloudWatch logs for queryable access.

WAF logs can be channeled via Kinesis Firehose to an S3 bucket.

Every KMS event is automatically captured in CloudTrail

| Attempt to query Cloudtrail for a known action (e.g. Deny or Approve Decryption) this must be retirieved within 15 minutes.

Attempt to delete a CloudTrail log, this should ideally fail.

Check that SHA-256 digest files are present and unbroken. | Yes ✅ | | Cost Efficiency | Ensure costs stay reasonable and remain below 2x the current cost. | Ensure that DR region runs at minimum viable cost (pilot-light).

Choice of Route 53 over Global Accelerator saves approx. $50 per month.

Choice of 2 AZs instead of 3 saves approx. $130 per month. | Pilot-Light DR: The Always-On option costs ~$180 per month at a baseline and offers a better cost alternative than the Active Failover which adds approx. $200 extra.

Utilize Graviton2 instances which are eco-friendly and less expensive (savings of 20%).

S3 Intelligent-Tiering: This automatically reduces storage costs by tiering every data stored by requency of access. | Using AWS Cost Explorer to track actual multi-region cost as a percentage (%) of primary-only cost.

Target result should be less than 150% of primary.

Establish a monthly review cadence to keep track of costs.

Although the WAF adds an extra $25 to the monthly cost, it is justified. | Yes ✅

(simulation result should give approx. 115% of primary region) |

B. Architecture Options Considered

Three (3) fundamentally different viable approaches were considered and evaluated, and each was assessed against every business requirement.

| | Option A Multi-AZ (Single-Region) Verdict: Rejected ⛔⛔ | Option B Pilot Light (Active-Passive) Verdict: ✅✅ | Option C Pilot Light (Active-Active) Verdict: Rejected⛔⛔ | | --- | --- | --- | --- | | Recovery Time Objective (RTO) | 30 - 60 minutes | 15 - 17 minutes | Less than 2 minutes | | Recovery Point Objective (RPO) | Less than 30 seconds (via Aurora Multi-AZ) | Less than 1 second (via Aurora Global) | Zero seconds (dual writes enabled) | | Uptime | Approximately 99.9% (effective only within AZs and not cross-regional) | Approximately 99.95% (cross-regional) | Approximately 99.999% (cross-regional) | | Cost Implications | ~ 110% of current | ~ 115 - 120% of current | ~ 200% of current | | Architectural Complexity | Low | Medium | Very High | | Summary | This option fails the 15 minute RTO especially if the primary region of presence (af-south-1) experiences a regional outage, there is no secondary to failover to. The service remains down until AWS recovers the region and this could run into hours. This fails by a factor of 4 - 16X. | This option meets the requirements within the lowest justifiable cost. The DR region runs a baseline pilot light (active-passive). On failover, it easily scales to full capacity within the 15 minute window. The only accepted tradeoff here is that the RTO may extend slightly above 15 minutes dur to DNS propagation. This will have to be documented with the bank as an accepted risk. | This is technically a more superior option, however, it fails the cost requirement. Secondly, a dual write conflict complexity is introduced with this option making it an overkill vs. the current scale of operations. This option might be considerd as an upgrade plan if the transactions grow significantly above 10X and bank decides on hard requirements. |

C. Architecture Decision Records (ADRs)

ADR 01: Region Selection

Context

TradeCoreAfrica operates across Nigeria, Ghana, and Kenya, serving 340 active business clients with NGN 180 million in daily transactions. A Tier-1 Nigerian bank distribution partner requires formal compliance with Nigerian data sovereignty regulations. The bank contract mandates that all customer financial data must remain within Nigeria's geographic boundaries at all times. AWS currently offers one African region: af-south-1 (Cape Town, South Africa). A new region, af-west-1 (Lagos, Nigeria), is scheduled but not yet generally available as of this ADR.

Decision

Deploy the primary production workload to af-south-1 (Africa, Cape Town) as the closest available AWS region to Nigeria's client base, while actively monitoring AWS's af-west-1 (Lagos) launch timeline for future migration.

Alternatives Considered