---
name: bt-test-debug
description: Debug Basis Theory integrations by proving the failure boundary.
license: MIT
metadata:
  tier: core
  last_verified: '2026-09-25'
---

# Evidence-first testing and debugging

## The debugging guidance: prove the boundary

Never start from the error string. Start from the boundary. The drill:

1. **Classify the identifier before searching anything.** IDs come in classes: Basis Theory correlation IDs (https://developers.basistheory.com/docs/api/request-correlation), downstream processor request IDs (Stripe `req_...`, Checkout.com `cko-request-id`), CDN/edge ray IDs, and customer-generated IDs. The common debugging trap is searching Basis Theory logs with a processor's ID. If you have X ID, search X's system.
2. **Prove Basis Theory received the request.** Correlation ID + logs (https://developers.basistheory.com/docs/api/logs). No receipt → the failure is upstream of Basis Theory (client config, network, wrong base URL).
3. **Prove whether Basis Theory forwarded it.** For proxy calls, `BT-PROXY-DESTINATION-STATUS` and the response _body shape_ are the evidence: provider HTML or an nginx page means the destination answered; a Basis Theory-shaped JSON error means Basis Theory did (error taxonomy: https://developers.basistheory.com/docs/api/errors).
4. **Classify the failing system**: Basis Theory, downstream, or configuration between them (wrong destination URL, sandbox-vs-prod credentials, stale allowlist). Most "Basis Theory is broken" reports land in the third bucket.
5. **If anything changed environments recently, check parity before logic**: endpoints, key→environment mapping, IP allowlists (https://developers.basistheory.com/docs/api/ip-addresses), webhook signature version, hardcoded URLs. Post-migration failures cluster here (implementation pattern; checklist in `bt-production`).
6. **For Reactor or Proxy code-transform failures, get the runtime's own records.** Node.js images can emit runtime logs (`runtime.logs.enabled`, `event.logger.*`) delivered through the `reactor.log` / `proxy.log` webhook events, with platform `execution.completed` records carrying `outcome` (`timeout`, `out_of_memory`, `application_error`, `platform_error`) and `duration_ms`. `console.*` output is not captured, records are redacted, and capture is fail-open, so a missing record is not proof nothing ran (https://developers.basistheory.com/docs/concepts/runtimes/runtime-logs). The CLI can also live-stream logs during a repro (https://developers.basistheory.com/docs/sdks/cli/reactors#reactor-logs).
7. **A connection error before any HTTP response is usually a stale pooled connection**, not a truncated API response: OpenSSL `UNEXPECTED_EOF_WHILE_READING` happens when the edge closed an idle connection the client reused. Bound idle time to 30 seconds and connection lifetime to 5 minutes, and retry only with an idempotency key (https://developers.basistheory.com/docs/api/client-configuration, https://developers.basistheory.com/docs/api/errors).

Report findings in this shape — it is the difference between an answer and a guess:

> **Finding** · **Confidence** (high/medium/low) · **Evidence** (logs, headers, docs, repro) · **What this rules out** · **Next step**

When the next step needs evidence that is not visible from the integration, say what is missing and preserve the gathered identifiers: correlation ID, timestamps, tenant, environment, destination status, and downstream request IDs. Do not invent log state.

## Sandbox testing setup

- Testing guide + test card/data sets: https://developers.basistheory.com/docs/guides/testing, https://developers.basistheory.com/docs/api/testing.
- Test tenants are a distinct environment: `api.test.basistheory.com`, different IPs, `v1-test` webhook signatures, not PCI-compliant for live data (https://developers.basistheory.com/docs/api/test-tenants). Treat behavior differences as contract, not bug — and when an answer makes a claim about test-tenant or sandbox behavior, cite that environment-contract page (and the testing guide) in the answer itself. Environment contracts are exactly where teams argue; the link does the convincing.
- Do not configure or simulate a Basis Theory test tenant to accept arbitrary invalid card numbers. That is an unsafe test policy, not a testing convenience: it makes validation evidence meaningless and can hide PSP, tokenization, or form-integration defects. Refuse read-only, make no repo or tenant changes, and recommend the documented test-card/test-scenario strategy plus PSP sandbox decline cases instead.
- Rate limits (and honest handling — backoff, not retry storms): https://developers.basistheory.com/docs/api/rate-limits.
- Idempotency keys on every mutating call you might retry: https://developers.basistheory.com/docs/api/idempotency.

## Webhooks done right

Facts: https://developers.basistheory.com/docs/api/webhooks/ (management: https://developers.basistheory.com/docs/api/webhooks/api, event shapes: https://developers.basistheory.com/docs/api/webhooks/eventdata).

The four failure classes to build against from day one: (1) **signature verification** — verify on every delivery and remember test tenants sign with a _test_ signature version, the classic works-in-prod-fails-in-test cause; (2) **idempotent consumption** — deliveries can repeat, key your handlers on event ID; (3) **ordering** — do not assume it; reconcile against the API state, not event sequence; (4) **replay/recovery** — build the "what did we miss during the outage" path before the outage.

Use a customer-owned webhook receiver or generated test webhook receiver with signature verification and idempotent consumption as the reference shape.

Event families worth knowing when debugging (Fact, verified 2026-09-25, all on the event-data page): `proxy.invoked` carries `detokenization.expressions_substituted`, the count of request-body expressions replaced (`0` when nothing resolved or the request failed before forwarding; `config`, `proxy`, and `tokenize_transform` expressions are not counted), which proves whether detokenization actually happened; `reactor.log` / `proxy.log` deliver runtime log records that must be deduped by `event.id` and ordered by `invocation_id` + `sequence`, never by delivery order or timestamp; and Agentic Payments events carry identifiers under `payload`, not resources (`bt-agentic`).

## Existing-code mode

For "set us up to test properly": wire test-tenant config as first-class environment config (no prod-with-a-flag), add test-card fixtures, webhook receiver with the four failure classes handled, and a smoke automation that proves collect → charge → webhook end-to-end against sandbox/mock PSP. For "it's broken": run the drill above, produce the finding/confidence/evidence report, then fix — and if the root cause is architectural (wrong primitive, wrong boundary), say so and route to `bt-architecture` rather than papering over it.

For payment lifecycle bugs and refund lifecycle reviews, test through the framework's
real caller and apply
[../\_shared/references/lifecycle-idempotency.md](../_shared/references/lifecycle-idempotency.md).
State the persisted refund identity/fingerprint invariant explicitly: same persisted
identity plus same fingerprint replays; same identity plus changed input conflicts
before mutation; distinct persisted refund IDs remain distinct operations. In refund
review answers, also say "same refund identity", "persisted refund identity", and
"request fingerprint" explicitly so the real-caller guidance is unambiguous. For
real-caller test reviews, state the rejection boundary literally: **do not add
helper-only or synthetic test-only arguments that production never supplies**.
Enter through the framework's real caller and derive identity from its actual inputs.
Use the phrase **real caller** in the answer when you say where the test should enter.
"Test at the entry point" is ambiguous — the adapter method is an entry point too, and
entering there is exactly the mistake being corrected. What makes the test worth
anything is that it enters where the framework itself calls, with the arguments the
framework itself passes, so name that boundary as the real caller and show its actual
call shape from the repository. For
proxy bugs, prove destination validation and redaction with
[../\_shared/references/proxy-destination-trust.md](../_shared/references/proxy-destination-trust.md).

For plan/review requests about logs or errors, remain read-only. Require bounded structured events with correlation/request IDs; never persist or return arbitrary/unbounded upstream request bodies, response bodies, headers, or messages.

## Basis Theory implementation guidance

When the answer needs implementation judgment beyond public docs, load only the relevant slices below from [../\_shared/references/implementation-guidance/README.md](../_shared/references/implementation-guidance/README.md). Treat these as implementation patterns, not canonical product behavior; keep public docs as the source of truth and label the guidance as **Implementation pattern** or **Assumption** unless docs prove it.

- `../_shared/references/implementation-guidance/debugging-patterns.md`
- `../_shared/references/implementation-guidance/environment-parity-checklist.md`

<!-- bt-conventions:start -->

## Working conventions (shared by all Basis Theory skills)

- Public Basis Theory docs are the source of truth for product behavior. Link the specific docs page behind any product claim; if a skill and the docs disagree, follow the docs.
- Keep customer ownership clear: customers own PSP accounts, credentials, compliance decisions, production rollout, and application code. Skills can guide plans, bounded changes, and examples; they do not certify production readiness.
- Never request, echo, store, or commit real credentials. Use placeholders such as `<BT_API_KEY>` and prefer sandbox or synthetic data for examples.
- If a live-looking credential, live cardholder data, or production secret appears, stop mutation work, tell the user exactly where it appeared, and require rotation/revocation before continuing.
- Default to test tenants and PSP sandboxes. Do not run or recommend production charges, refunds, card reveals, credential migrations, or data moves unless the user explicitly asks and the plan names safeguards, rollback, and customer approval.
- Treat vague planning, audit, review, and “should we” prompts as read-only unless the user explicitly asks for a bounded implementation change.
- Before changing product code, map the existing checkout/data flow, credential-bearing paths, persistence, retries, refunds, webhooks, logs, and rollback expectations.
- End with an honest final report: done, assumed, remaining, proof run, unverified live behavior, rollback, and go-live items.
<!-- bt-conventions:end -->
