Designing a Transactional Email System: Delivery Infrastructure, Template Rendering, and Deliverability Monitoring at Scale
How to architect transactional email for SaaS products at scale: provider selection and failover between SendGrid, Postmark, and AWS SES, MJML template rendering pipelines, delivery guarantees via queuing and retry, bounce and complaint handling, and reputation monitoring with TypeScript implementation patterns.
Most SaaS teams treat email as a footnote. They call sendgrid.send() somewhere in the codebase and ship it. This works until it does not: a burst of sign-ups causes rate limiting, a campaign triggers spam filters, bounce rates climb silently, and then one morning the welcome email stops arriving for half of new users. The deliverability debt comes due all at once.
Transactional email is infrastructure, not a utility call. This article covers how to architect it: provider selection and failover, template rendering with MJML, delivery guarantees through queuing and retry, bounce and complaint processing, and reputation monitoring. All patterns use TypeScript.
The Problem with a Single Provider
The simplest architecture is a direct API call to one provider. It is also the most brittle.
Providers fail. SendGrid had a major outage in 2024. Postmark occasionally rate-limits new accounts during reputation warm-up. AWS SES rejects messages when a sending identity has not been fully verified. None of these are edge cases; they are normal operational events in a multi-year product lifecycle.
The second problem is rate limits and warmup. When you migrate to a new IP pool or a new provider, inbox providers (Gmail, Outlook, Yahoo) apply heightened scrutiny. You cannot go from zero to 50,000 emails a day on a fresh IP. The volume must be ramped over two to four weeks. If you only have one provider and that provider’s shared IP pool gets blocklisted (which happens), you have no fallback.
A production email system needs at minimum two providers and a routing layer that can shift traffic between them.
Provider Abstraction Layer
Start with an interface that every provider adapter implements. This keeps the switching cost low and makes testing straightforward.
export interface EmailProvider {
name: string;
send(message: EmailMessage): Promise<EmailSendResult>;
isHealthy(): Promise<boolean>;
}
export interface EmailMessage {
to: string;
from: string;
subject: string;
html: string;
text: string;
messageId: string;
tags?: Record<string, string>;
}
export interface EmailSendResult {
provider: string;
providerMessageId: string;
accepted: boolean;
error?: string;
}
A SendGrid adapter using their Node.js client:
import sgMail from "@sendgrid/mail";
export class SendGridProvider implements EmailProvider {
name = "sendgrid";
constructor(private readonly apiKey: string) {
sgMail.setApiKey(apiKey);
}
async send(message: EmailMessage): Promise<EmailSendResult> {
const [response] = await sgMail.send({
to: message.to,
from: message.from,
subject: message.subject,
html: message.html,
text: message.text,
customArgs: { messageId: message.messageId, ...message.tags },
});
return {
provider: this.name,
providerMessageId: response.headers["x-message-id"] as string,
accepted: response.statusCode === 202,
};
}
async isHealthy(): Promise<boolean> {
// SendGrid does not expose a dedicated health endpoint.
// Track recent error rates instead of polling.
return true;
}
}
A Postmark adapter for a failover path:
import { ServerClient } from "postmark";
export class PostmarkProvider implements EmailProvider {
name = "postmark";
private readonly client: ServerClient;
constructor(serverToken: string) {
this.client = new ServerClient(serverToken);
}
async send(message: EmailMessage): Promise<EmailSendResult> {
const result = await this.client.sendEmail({
To: message.to,
From: message.from,
Subject: message.subject,
HtmlBody: message.html,
TextBody: message.text,
MessageStream: "outbound",
Metadata: { messageId: message.messageId },
});
return {
provider: this.name,
providerMessageId: result.MessageID,
accepted: result.ErrorCode === 0,
};
}
async isHealthy(): Promise<boolean> {
return true;
}
}
Failover Router
The router tries the primary provider and falls back to secondary on failure. It tracks consecutive errors so a degraded provider is deprioritized without requiring a manual intervention.
export class EmailRouter {
private primaryErrors = 0;
private readonly errorThreshold = 5;
constructor(
private readonly primary: EmailProvider,
private readonly fallback: EmailProvider
) {}
async send(message: EmailMessage): Promise<EmailSendResult> {
const usePrimary = this.primaryErrors < this.errorThreshold;
const provider = usePrimary ? this.primary : this.fallback;
try {
const result = await provider.send(message);
if (!result.accepted) {
throw new Error(`Provider ${provider.name} rejected the message`);
}
if (provider === this.primary) {
this.primaryErrors = Math.max(0, this.primaryErrors - 1);
}
return result;
} catch (error) {
if (provider === this.primary) {
this.primaryErrors += 1;
// Attempt fallback immediately
return this.fallback.send(message);
}
throw error;
}
}
}
For AWS SES as a third option, the pattern is identical: implement EmailProvider, inject the SES client, and add it to the router chain.
Template Rendering with MJML
HTML email is a mess. Every email client has its own quirks: Outlook renders using Microsoft Word’s layout engine, Gmail strips <style> tags from the <head>, Apple Mail is relatively compliant. Writing raw HTML that looks correct across all of them is tedious and error-prone.
MJML solves this by providing a component-based markup language that compiles to battle-tested HTML. The output uses inlined styles and table-based layouts that work across clients.
A rendering pipeline that takes a template name and data variables and returns the final HTML:
import mjml2html from "mjml";
import Handlebars from "handlebars";
import fs from "node:fs";
import path from "node:path";
interface RenderOptions {
template: string;
variables: Record<string, unknown>;
}
interface RenderedEmail {
html: string;
text: string;
subject: string;
}
export class TemplateRenderer {
private readonly templateDir: string;
private readonly cache = new Map<string, HandlebarsTemplateDelegate>();
constructor(templateDir: string) {
this.templateDir = templateDir;
}
async render(options: RenderOptions): Promise<RenderedEmail> {
const mjmlSource = this.loadTemplate(`${options.template}.mjml`);
const compiledMjml = this.compile(
`${options.template}.mjml`,
mjmlSource
)(options.variables);
const { html, errors } = mjml2html(compiledMjml, {
validationLevel: "strict",
});
if (errors.length > 0) {
throw new Error(
`MJML compilation errors: ${errors.map((e) => e.formattedMessage).join(", ")}`
);
}
const textSource = this.loadTemplate(`${options.template}.txt`);
const text = this.compile(
`${options.template}.txt`,
textSource
)(options.variables);
const subjectSource = this.loadTemplate(`${options.template}.subject.txt`);
const subject = this.compile(
`${options.template}.subject.txt`,
subjectSource
)(options.variables).trim();
return { html, text, subject };
}
private loadTemplate(filename: string): string {
return fs.readFileSync(
path.join(this.templateDir, filename),
"utf-8"
);
}
private compile(
key: string,
source: string
): HandlebarsTemplateDelegate {
if (!this.cache.has(key)) {
this.cache.set(key, Handlebars.compile(source));
}
return this.cache.get(key)!;
}
}
Keep one directory per template type: templates/welcome/, templates/password-reset/, templates/invoice/. Each directory contains the .mjml source, a .txt plaintext fallback, and a .subject.txt file. The plaintext version matters: some inboxes prioritize spam classification based on the text part, and some users read plain text by preference.
Delivery Guarantees via Queuing
Fire-and-forget email calls inside request handlers are a reliability anti-pattern. If the provider call fails, the email is silently lost. If it times out, it blocks the request. The solution is a queue.
The send path should write a job and return. The queue worker handles the actual delivery and retry logic.
export interface EmailJob {
id: string;
to: string;
template: string;
variables: Record<string, unknown>;
from: string;
replyTo?: string;
attempts: number;
scheduledAt: Date;
createdAt: Date;
}
export async function enqueueEmail(
db: DatabaseClient,
job: Omit<EmailJob, "id" | "attempts" | "scheduledAt" | "createdAt">
): Promise<string> {
const id = crypto.randomUUID();
await db.query(
`INSERT INTO email_jobs
(id, to_address, template, variables, from_address, reply_to, attempts, scheduled_at, created_at)
VALUES ($1, $2, $3, $4, $5, $6, 0, NOW(), NOW())`,
[id, job.to, job.template, JSON.stringify(job.variables), job.from, job.replyTo ?? null]
);
return id;
}
The worker processes jobs with exponential backoff on failure:
export class EmailWorker {
constructor(
private readonly db: DatabaseClient,
private readonly renderer: TemplateRenderer,
private readonly router: EmailRouter
) {}
async processBatch(batchSize = 25): Promise<void> {
const jobs = await this.db.query<EmailJob>(
`SELECT * FROM email_jobs
WHERE scheduled_at <= NOW()
AND status = 'pending'
AND attempts < 5
ORDER BY scheduled_at ASC
LIMIT $1
FOR UPDATE SKIP LOCKED`,
[batchSize]
);
await Promise.allSettled(jobs.rows.map((job) => this.processJob(job)));
}
private async processJob(job: EmailJob): Promise<void> {
await this.db.query(
`UPDATE email_jobs SET attempts = attempts + 1, status = 'processing' WHERE id = $1`,
[job.id]
);
try {
const rendered = await this.renderer.render({
template: job.template,
variables: job.variables,
});
await this.router.send({
to: job.to,
from: job.from,
subject: rendered.subject,
html: rendered.html,
text: rendered.text,
messageId: job.id,
});
await this.db.query(
`UPDATE email_jobs SET status = 'sent', sent_at = NOW() WHERE id = $1`,
[job.id]
);
} catch (error) {
const nextAttempt = job.attempts + 1;
const backoffSeconds = Math.pow(2, nextAttempt) * 30; // 60s, 120s, 240s...
const scheduledAt = new Date(Date.now() + backoffSeconds * 1000);
await this.db.query(
`UPDATE email_jobs
SET status = 'pending', scheduled_at = $1, last_error = $2
WHERE id = $3`,
[scheduledAt, (error as Error).message, job.id]
);
}
}
}
FOR UPDATE SKIP LOCKED ensures that multiple worker instances can run concurrently without processing the same job twice. This is the same pattern used for reliable background job processing in any queue built on Postgres.
Bounce and Complaint Handling
Every email provider sends webhook events for delivery outcomes: bounces, spam complaints, opens, and clicks. Processing these correctly is non-negotiable.
Hard bounces mean the address does not exist or the receiving server permanently rejected the message. If you keep sending to hard-bounced addresses, your sender reputation degrades. Soft bounces are temporary failures (mailbox full, server temporarily unavailable) and should be retried.
Spam complaints mean the recipient clicked “This is spam” in their email client. Above 0.08% complaint rate, Gmail and Yahoo will start filtering your email to spam folders. Above 0.3%, they will start rejecting it outright.
export type BounceType = "hard" | "soft";
export type WebhookEventType = "bounce" | "complaint" | "delivered" | "open" | "click";
interface EmailEvent {
messageId: string;
eventType: WebhookEventType;
bounceType?: BounceType;
email: string;
timestamp: Date;
providerData: Record<string, unknown>;
}
export class EmailEventProcessor {
constructor(private readonly db: DatabaseClient) {}
async process(event: EmailEvent): Promise<void> {
await this.db.query(
`INSERT INTO email_events (message_id, event_type, email, timestamp, provider_data)
VALUES ($1, $2, $3, $4, $5)
ON CONFLICT (message_id, event_type) DO NOTHING`,
[event.messageId, event.eventType, event.email, event.timestamp, JSON.stringify(event.providerData)]
);
if (event.eventType === "bounce" && event.bounceType === "hard") {
await this.suppressEmail(event.email, "hard_bounce");
}
if (event.eventType === "complaint") {
await this.suppressEmail(event.email, "complaint");
}
}
private async suppressEmail(email: string, reason: string): Promise<void> {
await this.db.query(
`INSERT INTO email_suppressions (email, reason, created_at)
VALUES ($1, $2, NOW())
ON CONFLICT (email) DO UPDATE SET reason = EXCLUDED.reason, created_at = NOW()`,
[email.toLowerCase(), reason]
);
}
}
Before enqueuing any email, check the suppression list:
export async function isSuppressed(
db: DatabaseClient,
email: string
): Promise<boolean> {
const result = await db.query(
`SELECT 1 FROM email_suppressions WHERE email = $1`,
[email.toLowerCase()]
);
return result.rowCount > 0;
}
Each provider has a different webhook format and signature verification mechanism. Build a webhook ingestion layer that normalizes events into the common EmailEvent type. Postmark uses a JSON payload with a custom header. SendGrid uses an array of event objects signed with a verification key. Parse these at the HTTP layer before the event reaches the processor.
Reputation Monitoring
Deliverability problems are much easier to fix if caught early. The signals to monitor are:
Bounce rate. Track hard and soft bounces as a rolling percentage of messages sent. Hard bounce rate above 2% is a serious problem. Anything above 5% often triggers provider account review.
Complaint rate. Track spam complaints per messages sent. The industry threshold is 0.08% for Gmail’s Postmaster Tools. You should alert well below that.
Delivery latency. The time between send and delivery event from the provider webhook. Sudden spikes in delivery latency often precede deliverability issues.
Domain reputation. Google’s Postmaster Tools exposes domain reputation (high, medium, low, bad) and IP reputation for authenticated senders. Integrate with the Postmaster Tools API to pull these signals into your monitoring dashboard.
export interface DeliverabilityMetrics {
periodStart: Date;
periodEnd: Date;
sent: number;
delivered: number;
hardBounces: number;
softBounces: number;
complaints: number;
hardBounceRate: number;
complaintRate: number;
deliveryRate: number;
}
export async function computeMetrics(
db: DatabaseClient,
start: Date,
end: Date
): Promise<DeliverabilityMetrics> {
const result = await db.query<{
sent: string;
delivered: string;
hard_bounces: string;
soft_bounces: string;
complaints: string;
}>(
`SELECT
COUNT(*) FILTER (WHERE status = 'sent') AS sent,
COUNT(*) FILTER (WHERE event_type = 'delivered') AS delivered,
COUNT(*) FILTER (WHERE event_type = 'bounce' AND bounce_type = 'hard') AS hard_bounces,
COUNT(*) FILTER (WHERE event_type = 'bounce' AND bounce_type = 'soft') AS soft_bounces,
COUNT(*) FILTER (WHERE event_type = 'complaint') AS complaints
FROM email_jobs j
LEFT JOIN email_events e ON j.id = e.message_id
WHERE j.created_at BETWEEN $1 AND $2`,
[start, end]
);
const row = result.rows[0];
const sent = parseInt(row.sent, 10);
const hardBounces = parseInt(row.hard_bounces, 10);
const complaints = parseInt(row.complaints, 10);
const delivered = parseInt(row.delivered, 10);
return {
periodStart: start,
periodEnd: end,
sent,
delivered,
hardBounces,
softBounces: parseInt(row.soft_bounces, 10),
complaints,
hardBounceRate: sent > 0 ? hardBounces / sent : 0,
complaintRate: sent > 0 ? complaints / sent : 0,
deliveryRate: sent > 0 ? delivered / sent : 0,
};
}
Set alerts at 1% hard bounce rate and 0.05% complaint rate. That gives you a buffer before reaching the thresholds that trigger provider action.
Authentication: SPF, DKIM, and DMARC
No deliverability discussion is complete without mentioning authentication. Gmail and Yahoo both require SPF, DKIM, and DMARC for bulk senders. Without these, inbox providers classify your messages as suspicious regardless of content quality.
SPF lists the IP ranges authorized to send email on behalf of your domain. DKIM adds a cryptographic signature to each message that receiving servers verify against a public key published in your DNS. DMARC ties them together and tells receiving servers what to do when a message fails: reject it, quarantine it, or allow it (while sending you a report).
The configuration lives in DNS. Your email provider will give you the exact records to add. The operational work is making sure those records stay correct when you change providers or add a new sending domain.
DMARC reporting is worth enabling from day one. Providers send aggregate reports (RUA records) that show which IP ranges are sending email that claims to be from your domain. This surfaces third-party services sending on your behalf that you may have forgotten about, and it catches spoofing attempts.
Tradeoffs Table
| Decision | Option A | Option B | Guidance |
|---|---|---|---|
| Provider strategy | Single provider | Multi-provider with failover | Multi-provider after meaningful volume; single is fine at < 1K/day |
| Template rendering | Server-side per send | Pre-compiled static templates | Pre-compile for speed; per-send for high personalization |
| Queue backend | Postgres-based jobs | Redis/BullMQ | Postgres is simpler and sufficient up to ~100K sends/day |
| Retry strategy | Fixed interval | Exponential backoff with jitter | Backoff always; avoid thundering herd on provider recovery |
| Bounce handling | Suppress on hard bounce | Retry hard bounces | Never retry hard bounces; they permanently damage reputation |
| Event storage | Log to file | Append-only events table | Events table; required for dispute resolution and metrics |
Production Considerations
Warm up new IPs gradually. If you add a provider or migrate to a dedicated IP, start at 200 emails per day and double weekly. Inbox providers use volume ramp-up as a signal of legitimate senders. Jumping to full volume immediately triggers filtering.
Segment your sending streams. Transactional email (receipts, password resets, notifications) and marketing email should use separate IP pools and separate subdomains. A marketing campaign that damages reputation should not affect transactional delivery. Postmark enforces this separation by design. With SendGrid or SES, you configure it explicitly.
Test rendering across clients. Litmus and Email on Acid render your HTML in 90+ email clients. Run these checks in CI for any template change. A pixel misalignment in Outlook is not worth discovering after sending to 50,000 users.
Include unsubscribe headers. RFC 8058 and the List-Unsubscribe-Post header support one-click unsubscribe in Gmail and Outlook. Gmail requires this for bulk senders. Add it to every non-transactional email and handle the resulting POST requests to update your suppression list.
Monitor Postmaster Tools. Google’s Postmaster Tools is free and shows your domain reputation, IP reputation, delivery errors, and spam rates as seen by Gmail. Set up a recurring job that pulls these metrics into your observability stack so you see trends before they become incidents.
Set provider-level suppression lists. All three major providers (SendGrid, Postmark, SES) maintain suppression lists that prevent sending to bounced or unsubscribed addresses at the API layer. Keep your application-level suppression list in sync with the provider list. If you switch providers, export and re-import the suppression list before sending any volume.
Closing
The systems that fail at scale are usually the ones that were designed for the present load only. Email is one of them. A direct API call works at 100 sends per day. It becomes a liability at 100,000. Build the queue, the suppression list, the bounce processor, and the metrics from the beginning. The marginal cost of doing it right early is low. The cost of retrofitting it after a deliverability collapse, complete with re-warming IPs and repairing sender reputation, is much higher.
More in System Design
How Amazon Aurora Works Internally: The Log-Is-the-Database Architecture, Quorum Writes, and Storage-Compute Separation Behind Cloud-Native SQL
A deep dive into Aurora's storage-compute separation, the log-is-the-database design that pushes redo log processing to storage nodes, quorum-based replication across 6 copies in 3 AZs, protection group architecture, fast cloning, Serverless v2 scaling mechanics, and honest tradeoffs versus RDS, self-managed PostgreSQL, CockroachDB, and Cloud Spanner.
How CockroachDB Works Internally: Distributed SQL, Raft Consensus Per Range, and the Architecture Behind Serializable Transactions at Global Scale
A deep dive into CockroachDB's internals: the 512MB range-based data model with automatic splitting, per-range Raft replication, MVCC timestamp ordering with hybrid logical clocks, DistSQL physical planning, serializable transactions with timestamp refreshes, closed timestamps for follower reads, online schema changes using multi-version schema descriptors, and production considerations for hotspots and write amplification.
How Container Runtimes Work Internally: Namespaces, Cgroups, and the OCI Stack From Docker Run to Process Isolation
Containers are not virtual machines and they are not magic. They are a thin composition of Linux kernel primitives: namespaces, cgroups, and a layered filesystem. This article traces exactly what happens between docker run and a running process, covering the OCI spec, containerd, runc, overlay filesystems, and container networking at the kernel level.
How YugabyteDB Works Internally: DocDB Storage, Tablet Splitting, and the Dual API Architecture That Scales PostgreSQL Horizontally
A deep dive into YugabyteDB internals covering the DocDB storage engine built on RocksDB LSM trees, per-tablet Raft consensus, YSQL and YCQL dual query layers, distributed MVCC with hybrid logical clocks, automatic tablet splitting and rebalancing, xCluster multi-region replication, and production considerations for hotspots, connection pooling, and schema design.