AWS Engineer Interview Questions: Full Hiring Guide
Last updated: August 26, 2026
A lot of AWS engineer interviews spend too much time on questions a search engine could answer and not enough time on questions that reveal judgment. A candidate reciting the difference between an Application Load Balancer and a Network Load Balancer tells you they studied. A candidate walking through why they’d choose one over the other for your specific traffic pattern tells you they can actually do the job.
This guide is organized by interview stage, with the specific questions we use when screening nearshore AWS engineers for US clients, what a strong answer sounds like, and what to watch for as a red flag.
Core AWS and Networking Questions
Start here to confirm foundational depth before moving into scenario-based questions.
- Walk me through how you’d design a VPC for a three-tier application with public and private subnets across two availability zones.
- What’s the difference between a security group and a network ACL, and when would you use both together?
- How does an Application Load Balancer differ from a Network Load Balancer, and which would you choose for a WebSocket-heavy application?
- Explain how Auto Scaling Groups decide when to add or remove instances, and what can go wrong with a poorly tuned scaling policy.
- What’s the difference between S3 storage classes, and how would you design a lifecycle policy for logs that need to be queryable for 30 days and retained for compliance for 7 years?
IAM and Security Questions
This is where a lot of resumes fall apart. Anyone can say “IAM” on a resume. Fewer people can articulate least-privilege design under questioning.
- How would you design an IAM role for a Lambda function that needs to read from one specific S3 bucket and write to one specific DynamoDB table, and nothing else?
- What’s the difference between an IAM role and an IAM user, and when should you never use long-lived access keys?
- How would you set up cross-account access between a staging and production AWS account without granting excessive permissions?
- What would you check first if GuardDuty flagged unusual API activity from an EC2 instance?
- How do you approach a security audit finding that a production IAM policy is too permissive, without breaking the service that depends on it?
Infrastructure-as-Code Assessment Guide
What to ask in a 60-minute live IaC session
Give the candidate a realistic, scoped task: write a Terraform module for an RDS instance with automated backups, or review an existing module and identify what’s wrong with it. Use a shared screen, not a whiteboard. You’re testing how they actually work, not how they perform under artificial pressure.
What to watch during the session
Do they think about state management and what happens if two engineers apply changes simultaneously? Watch whether they parameterize values that should be configurable instead of hardcoding them. It also matters whether they consider what happens on a failed apply, and whether they mention testing the plan output before applying. These habits separate someone who’s written a few Terraform files from someone who’s maintained infrastructure-as-code at scale.
Architecture and Reliability Questions
Give a scenario that matches your actual environment and ask them to reason through it out loud.
- Our AWS bill grew 40% in three months with no corresponding growth in traffic. How would you investigate that?
- Walk me through how you’d design a disaster-recovery strategy for a database that can tolerate 15 minutes of data loss but needs to be back online within an hour.
- A service is intermittently timing out under load. Walk me through your troubleshooting approach, starting from the alert.
- How would you approach migrating a monolithic application running on a single large EC2 instance to a containerized architecture without a big-bang cutover?
- What’s your process for running a Well-Architected Framework review, and what’s the most common finding you’ve seen?
Communication and Nearshore Fit Questions
Technical skill without communication discipline is a liability in a distributed infrastructure role, where a misunderstood decision during an incident can extend an outage.
- Tell me about a production incident you handled. Walk me through what happened, how you communicated it, and what changed afterward.
- Describe a time you disagreed with a team lead’s architecture decision. How did you raise it?
- How do you handle a situation where you’re unsure whether a change is safe to make, and the person who could confirm it is offline?
- Send a short written summary of your approach to [scenario] within 24 hours. (This should be a real async exercise, not just a spoken question.)
Good Answers vs. Red Flags
IAM design question
Good: Describes scoping the role to specific resource ARNs and specific actions, mentions using conditions where appropriate, and explains why broad wildcard permissions are a risk even when convenient. Red flag: Defaults to attaching a managed AdministratorAccess policy “to keep things simple” or can’t explain the difference between a role and long-lived credentials.
Cost investigation question
Good: Starts with Cost Explorer or Cost and Usage Reports, breaks down spend by service and tag, and mentions checking for orphaned resources like unattached EBS volumes or idle NAT Gateways. Red flag: Jumps straight to “buy Reserved Instances” without diagnosing what’s actually driving the increase.
Incident question
Good: Describes a clear timeline, what they communicated and when, what the actual root cause was, and a specific follow-up action that prevented recurrence. Red flag: Vague on timeline, can’t explain root cause beyond “it just started working again,” or no mention of any follow-up.
Written communication exercise
Good: Clear structure, specific technical detail, appropriate length, no ambiguity about what they’re proposing. Red flag: Vague, overly long with no structure, or technical claims that don’t hold up to a follow-up question.
Frequently Asked Questions
Interview Format and Structure
How long should an AWS engineer technical interview take?
Plan for a 30-minute phone screen, a 60 to 90 minute live infrastructure-as-code or architecture session, and a separate 20 to 30 minute communication and culture conversation. Total technical evaluation time typically runs 2 to 2.5 hours across two or three separate sessions, not one marathon interview.
Should I use a take-home assignment or live coding for AWS engineers?
Live is stronger for infrastructure roles. A take-home assignment doesn’t reveal how a candidate thinks out loud, handles ambiguity, or responds to a follow-up question that changes the requirements mid-task, all of which matter more for infrastructure work than for a solo coding exercise.
Evaluating Certifications and Common Mistakes
How do you evaluate an AWS engineer’s certifications during an interview?
Treat certifications as a starting filter, not a scoring criterion. Ask the candidate to explain a concept from the certification in the context of a real scenario rather than asking them to define it abstractly. A candidate who can only recite the textbook definition of a concept but can’t apply it to your environment has a certification gap, not a skills gap.
What’s the biggest mistake hiring managers make interviewing AWS engineers?
Testing memorized service trivia instead of judgment. Questions like “what’s the max size of an S3 object” test recall that’s one search away from being answered. Questions like “how would you design access for this scenario” test the actual skill you’re hiring for. Weight the interview toward the second type.
Disclosure: Kore BPO is a nearshore and offshore staffing agency. This question set reflects our direct experience screening AWS engineers from Costa Rica for US infrastructure teams.
Skip the Interviews and Get Pre-Screened Candidates
Kore BPO already runs this exact screening process on every nearshore AWS engineer we place. Profiles in 2 to 5 business days, $0 upfront fees.
View Nearshore AWS Engineers


