A timeout is evidence about a dependency, not proof of a refund.
At 14:12, support sees three refunds for invoices whose owners say they did not request them. The policy service had elevated latency in the same window. That correlation is worth following, but the incident could also come from an incorrect policy result, a duplicate job, a replayed request, or a report that combines separate events.
Start with request IDs, policy-decision outcomes, refund ledger entries, and the payment provider’s operation IDs. Compare timestamps and invoice ownership without copying tokens or full payment details into the incident record. Use a disposable account and a fake refund writer for any reproduction.
- Asset
- Customer funds and the audit trail for billing changes.
- Caller controls
- The refund request and selected invoice; neither proves permission.
- Decision source
- A policy service decides whether this actor may refund this invoice.
- Invariant
- A refund is issued only after an explicit allow decision for that actor and invoice.
Input identifies a requested operation.
Allow, deny, or unavailable.
Money movement begins here.
Confirm what actually changed.
Fail open means a failure or missing rule permits the protected operation. Fail closed means the operation is denied or held when the required decision cannot be established. For a refund authorization check, “unknown” cannot safely be translated to “allowed.” That does not mean every system error should stop every feature; the boundary here is the protected side effect.
“Denied” and “could not ask” are different results.
A policy lookup has at least three outcomes: allow, deny, and unavailable or invalid. A boolean can represent the first two, but it cannot also explain why no decision arrived. Exceptions, Go errors, timeouts, malformed responses, and missing configuration all need a deliberate path.
Hold the actor, invoice, and requested action constant while you inspect the policy request and the refund operation. If the policy dependency returned an explicit deny, the caller should not proceed. If it timed out, the caller should also not proceed, but the user-facing response and operational alert may differ. Check whether a refund was sent to the provider, not just whether the API returned success.
Compare your diagnosis with the case fileReveal after naming a competing cause
- Observation
- Three refund records appear during a policy-service incident.
- Competing causes
- A catch-all fallback allowed the operation; the policy returned an incorrect allow; a retry duplicated an earlier refund; or ledger/provider records were correlated incorrectly.
- Discriminating check
- In a fixture, make the policy throw or return an error while holding the actor and invoice fixed. Record whether the refund writer is called. In the incident, reconcile request IDs and provider operation IDs against policy logs.
- Conclusion if confirmed
- If an unavailable policy result is followed by a refund writer call, the exception path violated the authorization invariant. That establishes this path; it does not by itself establish how many users or funds were affected.
The dangerous line is the fallback that invents an approval.
The examples keep the policy and refund writer behind small interfaces. They are isolated
functions, not live endpoints. The intentionally flawed branch records a dependency
failure, then sets allowed to true. The log line may help an operator, but it does
not undo the permission granted by the next line.
The languages express a failed dependency differently; the invariant is the same.
// Intentionally flawed: a policy lookup failure becomes permission to continue.
export async function issueRefundVulnerable(
policy: RefundPolicy,
refunds: RefundWriter,
logger: SafeLogger,
requestID: string,
actorID: string,
invoiceID: string
): Promise<void> {
let allowed = false;
try {
allowed = await policy.mayRefund(actorID, invoiceID);
} catch {
logger.warn({ requestID, event: 'refund_policy_unavailable' });
allowed = true;
}
if (!allowed) throw new RefundDenied();
await refunds.issue(invoiceID);
} // Intentionally flawed: a policy lookup failure becomes permission to continue.
func IssueRefundVulnerable(ctx context.Context, policy RefundPolicy, refunds RefundWriter, logger SafeLogger, requestID, actorID, invoiceID string) error {
allowed, err := policy.MayRefund(ctx, actorID, invoiceID)
if err != nil {
logger.Warn(requestID, "refund_policy_unavailable")
allowed = true
}
if !allowed {
return ErrRefundDenied
}
return refunds.Issue(ctx, invoiceID)
} Preserve three outcomes until the side effect is authorized.
The repair treats a dependency failure as a failed operation, not a denial and not an approval. An explicit denial remains a permission result; a timeout becomes a controlled unavailable result. The refund writer runs only after the policy returns an explicit allow.
At the HTTP boundary, map these internal outcomes deliberately. A denied request can return the product’s stable forbidden response. A policy outage can return a temporary service error or create a pending review item, depending on the product’s contract. Neither response should contain stack traces, policy payloads, or raw dependency credentials.
export async function issueRefundSafely(
policy: RefundPolicy,
refunds: RefundWriter,
logger: SafeLogger,
requestID: string,
actorID: string,
invoiceID: string
): Promise<void> {
let allowed: boolean;
try {
allowed = await policy.mayRefund(actorID, invoiceID);
} catch {
logger.warn({ requestID, event: 'refund_policy_unavailable' });
throw new PolicyUnavailable();
}
if (!allowed) throw new RefundDenied();
await refunds.issue(invoiceID);
} func IssueRefundSafely(ctx context.Context, policy RefundPolicy, refunds RefundWriter, logger SafeLogger, requestID, actorID, invoiceID string) error {
allowed, err := policy.MayRefund(ctx, actorID, invoiceID)
if err != nil {
logger.Warn(requestID, "refund_policy_unavailable")
return ErrPolicyUnavailable
}
if !allowed {
return ErrRefundDenied
}
return refunds.Issue(ctx, invoiceID)
} Review the complete isolated examplesInterfaces and both control paths
export interface RefundPolicy {
mayRefund(actorID: string, invoiceID: string): Promise<boolean>;
}
export interface RefundWriter {
issue(invoiceID: string): Promise<void>;
}
export interface SafeLogger {
warn(fields: { requestID: string; event: string }): void;
}
export class RefundDenied extends Error {}
export class PolicyUnavailable extends Error {}
// Intentionally flawed: a policy lookup failure becomes permission to continue.
export async function issueRefundVulnerable(
policy: RefundPolicy,
refunds: RefundWriter,
logger: SafeLogger,
requestID: string,
actorID: string,
invoiceID: string
): Promise<void> {
let allowed = false;
try {
allowed = await policy.mayRefund(actorID, invoiceID);
} catch {
logger.warn({ requestID, event: 'refund_policy_unavailable' });
allowed = true;
}
if (!allowed) throw new RefundDenied();
await refunds.issue(invoiceID);
}
export async function issueRefundSafely(
policy: RefundPolicy,
refunds: RefundWriter,
logger: SafeLogger,
requestID: string,
actorID: string,
invoiceID: string
): Promise<void> {
let allowed: boolean;
try {
allowed = await policy.mayRefund(actorID, invoiceID);
} catch {
logger.warn({ requestID, event: 'refund_policy_unavailable' });
throw new PolicyUnavailable();
}
if (!allowed) throw new RefundDenied();
await refunds.issue(invoiceID);
}
package secureerrors
import (
"context"
"errors"
)
var ErrRefundDenied = errors.New("refund denied")
var ErrPolicyUnavailable = errors.New("refund policy unavailable")
type RefundPolicy interface {
MayRefund(ctx context.Context, actorID, invoiceID string) (bool, error)
}
type RefundWriter interface {
Issue(ctx context.Context, invoiceID string) error
}
type SafeLogger interface {
Warn(requestID, event string)
}
// Intentionally flawed: a policy lookup failure becomes permission to continue.
func IssueRefundVulnerable(ctx context.Context, policy RefundPolicy, refunds RefundWriter, logger SafeLogger, requestID, actorID, invoiceID string) error {
allowed, err := policy.MayRefund(ctx, actorID, invoiceID)
if err != nil {
logger.Warn(requestID, "refund_policy_unavailable")
allowed = true
}
if !allowed {
return ErrRefundDenied
}
return refunds.Issue(ctx, invoiceID)
}
func IssueRefundSafely(ctx context.Context, policy RefundPolicy, refunds RefundWriter, logger SafeLogger, requestID, actorID, invoiceID string) error {
allowed, err := policy.MayRefund(ctx, actorID, invoiceID)
if err != nil {
logger.Warn(requestID, "refund_policy_unavailable")
return ErrPolicyUnavailable
}
if !allowed {
return ErrRefundDenied
}
return refunds.Issue(ctx, invoiceID)
}
Prove the operation does not happen on uncertainty.
Test the policy boundary with a fake refund writer that records calls. Cover an explicit allow, an explicit deny, a thrown exception or Go error, and an invalid or missing decision if the policy API can produce one. Assert the side effect count and the returned outcome. A generic “request failed” assertion can pass even if the refund already happened.
For a deployed incident, compare application decision logs with the payment provider’s operation history and ledger entries. A 500 response is not evidence that an external side effect rolled back. Also check retry behavior, because a second request can duplicate an operation that the first request completed before its response failed.
Exactly one refund is issued.
No refund writer call occurs.
No refund writer call occurs; caller receives controlled failure.
Idempotency prevents a duplicate provider operation.
Choose a response that protects the operation and helps recovery.
The policy service times out during a refund request. What should happen?
The actor is authenticated, but the policy service did not return a decision. The payment provider has not yet been called in this request path. Which next action preserves the security property and gives the user a recoverable result?
The confirmed bug in the example is narrow: policy unavailability becomes an allow decision and the refund writer runs. Repeated exploitation could increase financial impact if the caller can trigger refunds for invoices they do not control and the payment integration honors each request. The actual exposure depends on the endpoint’s reachable callers, policy rules, refund limits, provider behavior, and idempotency controls.
For guidance, see OWASP’s Authorization Cheat Sheet section on safely exiting failed access-control checks, and the Go documentation on error handling for database transactions for a separate example of why errors must be handled before committing state. JavaScript’s promise guide describes rejection propagation and explicit catch behavior. Checked 2026-10-01; the references support the language and transaction behaviors, not this fictional billing policy.
Connections to follow nextRelated lessons
Fail open or fail closed develops the decision rule. Idempotency and replay resistance addresses duplicate operations after retries.