Serverless in Production: Lessons Learned After Two Years

After two years running serverless in production, we share real-world insights on performance, cost, cold starts, and tooling that every team should know.

Serverless in Production: Lessons Learned After Two Years

Is your company ready for AI? Download our free checklist →

Download checklist

Introduction

Two years ago, our team at Tanok Tech decided to go all-in on serverless architecture. We had heard the promises: no servers to manage, automatic scaling, pay-per-execution pricing, and faster time to market. After building and operating several production systems on AWS Lambda, Azure Functions, and Google Cloud Functions, we've gathered enough battle scars to share what really works—and what doesn't.

This post covers the critical lessons we learned about cold starts, observability, cost management, state management, and deployment strategies. Whether you're considering serverless or already using it, these insights will help you avoid common pitfalls.

The Cold Start Reality

Cold starts are the most talked-about issue in serverless, but their impact is nuanced. In the first year, we tried every mitigation: provisioned concurrency, keeping functions warm with pings, and optimizing runtime initialization.

What Worked

  • Provisioned concurrency for latency-sensitive endpoints. For example, our payment API runs with a minimum of 10 concurrent executions, eliminating cold starts entirely.
  • Choosing the right runtime. Python and Node.js initialise in under 100ms, while Java and .NET can take 1-2 seconds. We migrated a Java service to Node.js and saw p90 latency drop from 3s to 200ms.
  • Reducing bundle size. We shaved 40% off our Lamba package size by removing unused dependencies and using tree-shaking.

What Didn't

VPC-enabled Lambda functions with NAT gateways are notoriously slow. In one case, cold start times jumped to 10 seconds. We solved this by using AWS PrivateLink or simply moving non-sensitive data outside the VPC.

Observability is Harder Than Expected

Monitoring serverless functions is different from monitoring servers. You can't SSH into an instance; you have to rely on distributed tracing and logging.

Our Stack

  • AWS X-Ray for tracing requests across Lambda, API Gateway, and DynamoDB.
  • CloudWatch Logs with structured JSON logging for easier querying.

But we quickly hit limits: cold starts caused gaps in traces, and high-throughput functions generated so many logs that costs exploded. We now filter logs aggressively, keeping only errors and important business events.

Tip: Use Custom Metrics

We created custom metrics for business KPIs (e.g., order count, error rate per function) and sent them to CloudWatch with minimal cost. This gave us a dashboard that actually reflected user experience.

Cost Management: The Hidden Pitfalls

Serverless pricing seems simple: pay for what you use. But hidden costs can balloon your bill.

Want a personalized diagnostic? Complete our free checklist →

Download checklist

Unexpected Cost Drivers

  • API Gateway costs exceeded Lambda costs in our architecture. Each API call incurs a charge, and with millions of requests, it adds up.
  • Data transfer costs between services. We were calling DynamoDB from Lambda in the same region, but cross-AZ data transfer still appeared on the bill.
  • Provisioned concurrency is charged even when not in use, so we only enable it for critical paths.

Optimization Strategies

  • Use AWS Compute Optimizer to right-size memory allocations. Increasing memory doesn't just speed up execution; it also reduces duration, often lowering cost.
  • Implement caching with ElastiCache or CloudFront to reduce function invocations.
  • Enable AWS Budgets and set alerts for anomalies.

State Management: Think Stateless

Serverless functions are ephemeral—they can be terminated at any time. We learned not to store state in memory or local disk.

What We Do

  • Use DynamoDB for session state, with TTL for automatic cleanup.
  • Use S3 for file uploads and Lambda to process them asynchronously.
  • Avoid sticky sessions by using idempotent API designs.

Pitfall: Distributed Transactions

We attempted a transaction that involved multiple services (Lambda -> Step Functions -> DynamoDB -> SQS). When a step failed, we had to implement a saga pattern with compensation logic. It was complex but necessary for consistency.

Deployment and Infrastructure as Code

Manual deployments are a no-go. We embraced Infrastructure as Code (IaC) from day one.

Our Tooling

  • AWS CDK for defining infrastructure in TypeScript. It provides higher-level constructs than CloudFormation.
  • CI/CD pipelines with GitHub Actions that deploy to dev, staging, and prod environments separately.
  • Canary deployments using AWS CodeDeploy to gradually shift traffic. This caught a bug that only appeared under full load.

Code Example: AWS CDK Lambda Function

import * as lambda from 'aws-cdk-lib/aws-lambda';

const myFunction = new lambda.Function(this, 'MyFunction', {
  runtime: lambda.Runtime.NODEJS_18_X,
  handler: 'index.handler',
  code: lambda.Code.fromAsset('lambda'),
  memorySize: 512,
  timeout: cdk.Duration.seconds(10),
  reservedConcurrentExecutions: 10,
});

This simply defines a Lambda function with sensible defaults. The CDK handles creating the role, log group, and permissions.

When Serverless is Not the Right Fit

We also learned where serverless falls short:

  • Long-running processes (e.g., video transcoding) exceed Lambda's 15-minute timeout. Use Fargate or dedicated servers.
  • Low-latency, high-throughput scenarios like real-time gaming. The overhead of function invocation adds too much latency.
  • Predictable high loads where reserved instances are cheaper than per-invocation pricing.

Final Thoughts

Serverless has transformed how we build software, but it's not a panacea. After two years, our key takeaways are:

  1. Always plan for cold starts, especially at high scale.
  2. Invest in observability from day one—it's harder to add later.
  3. Monitor costs aggressively, as they can surprise you.
  4. Embrace stateless designs and IaC.
  5. Know when to use serverless and when to skip it.

Would we do it again? Absolutely. The agility and scalability outweigh the challenges—as long as you go in with eyes wide open.

---

Have you had similar experiences? Let us know in the comments or reach out on Twitter.

Ready for the next step? Evaluate your company with our free checklist →

Download checklist

Related posts