Every product runs on servers. The real question is whether your team should be the one managing them.
For most of the history of web software, the answer was yes. Someone provisioned virtual machines, patched operating systems, and guessed how much capacity next month's traffic would need. Serverless architecture changes that deal. You write the code, the cloud provider runs it, and you pay for what actually executes.
The pitch is compelling, and it's often right. But serverless architecture is not a better version of traditional infrastructure. It's a different set of tradeoffs. Teams that treat it as the default end up with slow APIs and surprise bills. Teams that dismiss it end up paying people to babysit servers that sit idle most of the day.
At Big Human, we've been building digital products for 15+ years, from early-stage MVPs to enterprise platforms, including Vine, which Twitter acquired shortly after launch. Our engineering teams design and deploy backends on AWS, Azure, and Google Cloud, and our custom API development work uses serverless patterns like AWS API Gateway and AWS Lambda where they fit and traditional servers or containers where they don't. We don't have a favorite. We have a process for choosing.
This guide explains what serverless architecture is, what you gain and give up by adopting it, and how to tell when it's the right call versus traditional infrastructure. If you're already past the research phase and want to talk through a specific project, we'd love to chat.
Quick summary
Serverless architecture hands servers, scaling, and runtime to the cloud provider, so your team only writes and deploys code and pays for what actually runs.
It's event-driven: API requests, file uploads, queue messages, and schedules trigger stateless functions, while managed services hold the data.
You gain less operational work, automatic scaling, and costs that track usage. You give up some latency consistency, portability, and low-level control.
Serverless fits APIs with variable traffic, background jobs, MVPs, and teams without dedicated DevOps. Steady, latency-critical, or long-running workloads usually fit traditional infrastructure better.
The best answer is decided workload by workload, and it's often a hybrid of functions, containers, and managed services.
Serverless architecture is a way of building and running applications in which the cloud provider manages servers, scaling, and runtime, while your team only writes and deploys code. The servers still exist. You just never see, configure, or maintain them.
In a serverless environment, your application is broken into functions or managed services that run in response to events. An event might be an HTTP request from a web app, a new file upload, a message landing in a queue, or a scheduled task that fires every night. The provider spins up compute to handle the event, runs your code, and shuts it down when the work is done.
Serverless computing is a subset of cloud computing, and it usually shows up in two forms:
Function-as-a-service (FaaS): You deploy small units of business logic as functions, and the provider runs each one on demand. AWS Lambda, Azure Functions, and Google Cloud Functions (now part of Google Cloud Run functions) are the best-known FaaS platforms.
Backend-as-a-service (BaaS): You use fully managed services for common backend needs instead of building them. Think managed databases, user authentication, file storage, and push notifications. Firebase, AWS Amplify, and Supabase are common BaaS options.
Most real serverless architectures combine both. FaaS handles custom business logic, and BaaS covers the commodity pieces nobody should rebuild from scratch.
Three traits define the model. There is no server management. There is automatic scaling, including down to zero when nothing is happening. And there is pay-as-you-go pricing, where you're billed for requests and execution time rather than idle capacity. Each of those traits is also where the tradeoffs come from.
Serverless architecture works by connecting event sources to stateless functions and managed services, with the cloud provider handling everything between the request and the response. A typical request moves through a few predictable layers.
Nothing runs until something happens. An event source can be an API call, a database change, a message in a queue, an item in a data stream, a file landing in storage, or a timer. This event-driven model is what lets serverless scale to zero. No events means no compute and no charge.
For web applications and mobile apps, most events start as HTTP requests. An API gateway sits in front of your functions. It receives the request, handles routing, and often manages rate limiting, request validation, and authorization. Then it invokes the right function. AWS API Gateway, Azure API Management, and Google Cloud API Gateway all fill this role. It's how teams build RESTful APIs without running a single web server.
Each function does one job. One might create a user record. Another might resize an image. When an event arrives, the provider finds or creates an execution environment, loads your code, and runs it. When traffic increases, the provider runs more copies in parallel. That parallelism is where serverless achieves scalability, and it's why a well-designed function remains scalable without anyone adding servers.
Functions are stateless. They don't remember anything between invocations. Any data that needs to persist is stored in a database, cache, or storage service.
Because functions are stateless, the surrounding managed services do the heavy lifting. Databases such as Amazon DynamoDB and Amazon Aurora Serverless store data. Queues like Amazon SQS buffer work so spikes don't overwhelm downstream systems. Event streams like Amazon Kinesis move high volumes of data between components. Identity services handle user authentication.
When a function hasn't run recently, the provider has to create a fresh execution environment before running your code. That setup delay is called a cold start. Once the environment is created, it stays warm for a while and handles subsequent requests much faster. How long a cold start takes depends on the runtime, package size, memory settings, and network configuration. Lightweight JavaScript and Python functions typically start faster than large Java or .NET functions.
The core difference between serverless and traditional infrastructure is who manages capacity and how you pay for it. With traditional infrastructure, your team runs servers that are always on and billed whether they're busy or idle. With serverless, the provider runs compute only when events arrive and bills you per use.
“Traditional infrastructure” covers a range of options. Most teams comparing their choices are looking at one of these:
Virtual machines: You rent servers, like Amazon EC2 instances, and own everything above the hardware: operating system updates, security patches, scaling rules, and uptime.
Containers and Kubernetes: You package the application with its dependencies and run it on a cluster. Kubernetes handles deployment, scaling, and recovery, but someone still manages the cluster, capacity, and upgrades. This takes real DevOps investment.
Platform services: Managed app platforms, such as AWS Elastic Beanstalk and Heroku, sit in the middle. They remove some server work but still bill for always-on instances.
Infrastructure management: Traditional infrastructure puts provisioning, patching, and capacity planning on your team. Serverless hands it all off to the cloud provider.
Scaling: Traditional setups scale via autoscaling rules you configure and pay for in advance of demand. Serverless scales automatically per request, down to zero.
Cost model: Traditional infrastructure charges for provisioned capacity 24/7. Serverless charges for requests and execution time.
Performance: Always-on servers respond consistently. Serverless can add latency on cold starts.
Control: Traditional infrastructure lets you tune the operating system, run long processes, and hold persistent connections. Serverless limits memory, execution time, and runtime customization.
Portability: Containers run almost anywhere. Serverless apps depend heavily on one provider's services.
Microservices is a design approach, not an infrastructure choice. It means splitting an application into small, independently deployable services. You can run microservices architectures on serverless functions, on Kubernetes, or on virtual machines. Microservices describe how you divide the system. Serverless and traditional infrastructure describe where each piece runs.
The line between the two worlds has also blurred. Services like AWS Fargate, Google Cloud Run, and Azure Container Apps run containers in a serverless way, so teams can keep container packaging while avoiding cluster management.
The tradeoff at the heart of serverless architecture is simple: you give up control and some predictability in exchange for less operational work and costs that track usage. Whether that's a good deal depends entirely on the workload.
No server management: Your team doesn't patch operating systems, rotate servers, or plan capacity. That time goes back into the product. For small development teams without a dedicated DevOps function, this is often the deciding factor.
Automatic scaling: A product that sees 10 requests per hour on Tuesday and 10,000 per minute during a launch doesn't need anyone to change a setting.
Pay-as-you-go pricing: For spiky or unpredictable workloads, paying per request often costs far less than keeping servers running for peak traffic.
Built-in high availability: Major providers run functions across multiple availability zones by default, which gives you redundancy that takes real effort to build yourself.
Faster time to market: Managed services for authentication, storage, and messaging enable teams to ship features rather than plumbing.
Consistent latency: Cold starts introduce delays for some requests. Providers offer fixes, such as provisioned concurrency on AWS Lambda and pre-warmed instances on Azure Functions, but those features bring back some of the always-on costs you were trying to avoid.
Freedom from vendor lock-in: Your functions might be portable, but the triggers, permissions, queues, and databases around them usually aren't. Moving from AWS Lambda to Azure Functions is rarely a lift-and-shift.
Simple debugging: One request might touch an API gateway, four functions, a queue, and two databases. Tracing bugs across that chain requires deliberate observability: structured logging, distributed tracing, and tools like AWS X-Ray, Datadog, or OpenTelemetry.
Unlimited run time: Functions have limits on memory, payload size, and execution time. Standard AWS Lambda functions time out at 15 minutes, for example. Providers have added options for longer-running workflows, but heavy batch jobs and persistent connections usually still fit better on traditional compute.
Predictable cost at scale: Per-invocation pricing is efficient for uneven traffic. For workloads that run hot around the clock, it can cost more than reserved servers.
Some concerns change shape rather than disappear. Security is one. The provider secures the underlying servers, but you still own permissions, API validation, and secrets management, and every function needs its own tightly scoped access. The architecture discipline is another. It's easy to create functions and hard to keep hundreds of them organized. Without clear conventions, a serverless system can become as tangled as any legacy monolith.
Serverless architecture is the right call for event-driven workloads with unpredictable or uneven traffic, where delivery speed matters more than low-level control. These are the scenarios where it tends to outperform traditional infrastructure.
RESTful APIs behind an API gateway are the classic serverless use case. Each endpoint maps to a function, scaling is automatic, and you only pay when requests come in. This works well for SaaS products, mobile app backends, and internal tools where traffic rises and falls through the day.
When a user uploads a photo or document, a storage event can trigger a function that resizes the image, scans it, or extracts text. The work runs in the background, scales with upload volume, and costs nothing when no one is uploading.
Nightly reports, data cleanup, subscription renewals, and reminder emails are natural fits. Instead of running a server all day to do a few minutes of work, a scheduled trigger runs a function exactly when needed.
Functions can consume data streams and queues to process clickstream data, IoT readings, payment events, or webhooks from third-party services. Queues absorb spikes, and functions work through the backlog at whatever rate downstream systems can handle.
For a new product with unknown demand, serverless keeps fixed infrastructure costs near zero and lets a small team ship quickly. If the product takes off, the architecture scales with it. If it doesn't, you haven't paid for idle servers. It's one reason serverless shows up so often in startup app development.
If nobody on the team wants to be on call for servers, serverless removes an entire category of operational risk. Small teams get production-grade availability without hiring for it.
Traditional infrastructure is the better choice when workloads are steady and heavy, latency must be consistent, or the system needs control and portability that serverless can't offer. Choosing servers or containers in these cases isn't old-fashioned. It's matching the tool to the job.
If a service handles a heavy load around the clock, reserved virtual machines or a well-tuned Kubernetes cluster usually cost less than paying per invocation. The pay-per-use advantage disappears when there's never any idle time.
Real-time trading, multiplayer features, live collaboration, and checkout flows with strict response targets can't tolerate occasional cold starts. Always-on compute gives you predictable performance without paying extra to keep functions warm.
Video encoding at scale, machine learning training, large batch jobs, and persistent WebSocket connections push against function limits. These workloads run better on containers or dedicated instances, where you control memory, hardware, and run time.
Organizations with multi-cloud mandates, on-premises components, or strict data residency rules often need infrastructure they can move and inspect. Containers offer that flexibility. Serverless ties you more closely to one cloud provider's managed services.
A large existing application rarely benefits from being broken into functions all at once. Rewriting a working monolith into hundreds of functions can add risk without adding value. Moving it to containers first, then carving off serverless pieces where they make sense, is usually safer.
In practice, the choice is rarely all-or-nothing. Many of the strongest architectures are hybrid: containers or virtual machines for the core, always-on services, and serverless functions for event handling, background jobs, and the spiky edges of the system. That's especially common in enterprise web apps, where different parts of the platform have very different traffic patterns.
You decide between serverless and traditional infrastructure by evaluating each workload individually, not by picking a single model for the entire product. We start with the workload rather than an infrastructure philosophy: we map what each part of the product needs to do, how often, and how fast, and then decide which pieces belong on functions, which belong on containers, and which should be a managed service. Choosing infrastructure before you understand the workload makes every later decision harder and more expensive to undo.
Every product is different, so the specifics shift from project to project, but the general approach typically looks like this.
1\. Map workloads and traffic patterns
List what the system does: user-facing requests, background jobs, integrations, reports. For each, estimate volume, how spiky it is, and how fast it needs to respond. This map shows which workloads benefit from scaling to zero and which will run hot all day.
2\. Flag the hard constraints
Identify anything that rules serverless in or out. That includes strict latency targets, processes that run longer than function limits, persistent connections, compliance or data residency rules, and any commitment to staying portable across cloud providers. Constraints settle some decisions before cost even comes up.
3\. Assess your team's operational capacity
Be honest about DevOps skills and bandwidth. A team with Kubernetes experience and on-call coverage can run traditional infrastructure well. A small product team is often better served by managed services. Runtime familiarity matters too. Node.js and Python are popular serverless choices because they start quickly and have mature tooling on every major platform.
4\. Model cost at today's traffic and at scale
Estimate the monthly cost for each workload on both models, at current traffic and at ten times current traffic. Include the costs people forget: data transfer, logging, provisioned concurrency, and the engineering time spent managing servers. The crossover point where one model becomes cheaper is often clearer than expected.
5\. Design boundaries that keep options open
Group functionality by business domain, like orders, payments, and notifications, and keep business logic separate from provider-specific code. Clear boundaries let you move a single workload from functions to containers, or the reverse, without rewriting the product.
6\. Build observability and CI/CD from day one
Whichever model you choose, set up structured logging, tracing, and alerts before the first production release. Define infrastructure as code using tools such as Terraform, AWS CDK, or the Serverless Framework, and integrate it into continuous integration and continuous delivery pipelines. This makes the architecture measurable, so future infrastructure decisions are based on real data.
Running this process across different kinds of products tends to produce a few recognizable patterns:
Early-stage product with unknown demand: Mostly serverless. Functions behind an API gateway, a managed database, and managed authentication keep costs low and let a small team focus on features.
Growing SaaS platform: Hybrid. The core application runs in containers to ensure predictable performance, while file processing, notifications, webhooks, and scheduled tasks run as functions.
High-traffic consumer app with real-time features: Mostly traditional infrastructure for the latency-sensitive core, with serverless reserved for background work and traffic spikes that don't touch the real-time path.
Enterprise system with legacy components: Containers first to modernize the core safely, then serverless for new integrations and event-driven workflows added around it.
None of these patterns are permanent. Traffic changes, teams grow, and provider pricing shifts. A good architecture makes it cheap to revisit the decision.
Resource allocation decisions, such as function memory, concurrency limits, and instance sizes, also come out of this process. They should be revisited once real traffic data exists.
If you're working through these architecture decisions on a real product, this is exactly what our process is built for. Reach out to us to chat about your project.
Serverless costs scale with usage, while traditional infrastructure costs scale with provisioned capacity. That means serverless is often cheaper for new, small, or spiky workloads and can become more expensive for heavy, constant ones. The real comparison depends on your traffic, not a general rule.
Request volume and how evenly it's spread across the day
Function memory settings and execution time
API gateway requests and any provisioned concurrency
Data transfer, storage, and database throughput
Logging and monitoring volume, which is often underestimated
Instance or cluster size, provisioned for peak demand
Idle capacity during low-traffic hours
Engineering time for patching, upgrades, scaling, and on-call
Redundancy needed to reach high availability
Cost optimization habits matter for either model. On serverless, right-size function memory, batch small events from queues, and set log retention policies. On traditional infrastructure, use reserved capacity for steady baselines and autoscaling for peaks. On both, review spending by workload so that expensive pieces are visible early.
Ready to talk scope? Reach out to Big Human, and we'll walk through what you're building and scope it with you as part of the consultation process.