To capture an exception in Python, catch the narrowest expected exception with try and except, then recover, translate it, add safe context, or terminate the current unit of work. Use logging.exception() inside the handler when that boundary owns the failure and needs the traceback; otherwise let the exception propagate to a boundary that does.
Capturing is not the same as suppressing. An empty except block can hide corrupted state, incomplete work, and programming defects. A useful handler makes an explicit control-flow decision and preserves enough evidence to understand the failure without leaking secrets.
Catch exceptions only where a decision is possible
Keep the try block small and catch the exception types that the protected operation is expected to raise:
from pathlib import Path
class ConfigurationError(RuntimeError):
pass
def read_worker_count(path: Path) -> int:
try:
text = path.read_text(encoding="utf-8")
except OSError as exc:
raise ConfigurationError("worker configuration is unavailable") from exc
try:
count = int(text.strip())
except ValueError as exc:
raise ConfigurationError("worker count is not an integer") from exc
if count < 1:
raise ConfigurationError("worker count must be positive")
return count
Each handler translates one known implementation failure into an application-level exception. raise ... from exc records the original exception as the direct cause, so an internal traceback retains both layers. The Python exception tutorial documents this explicit chaining form for transformed exceptions.
The new messages deliberately avoid inserting the file contents, path supplied by an untrusted user, or str(exc). Exception strings and .args can contain a URL with credentials, a database value, a local path, a request fragment, or another secret. Treat them as potentially sensitive even when a built-in exception usually looks harmless.
Catch Exception only at a boundary that can consistently end or translate one work unit: a command entry point, request adapter, message handler, scheduler invocation, or worker loop. Internal business functions should normally catch specific failures and let unexpected defects propagate.
Do not normally catch BaseException. SystemExit, KeyboardInterrupt, and asyncio.CancelledError use that hierarchy to control program or task lifetime. Python's exception-handling guidance identifies Exception as the base for non-fatal exceptions and recommends allowing unexpected exceptions to propagate.
Avoid both of these patterns:
def bad_update_silently() -> None:
try:
update_inventory()
except:
pass
def bad_update_catches_shutdown() -> bool:
try:
update_inventory()
except BaseException:
return False
return True
The bare handler catches termination signals, and both versions discard the reason and pretend the operation completed. If the operation is optional, catch its documented exception, record a safe outcome, and make the fallback visible.
Use else, finally, and context managers for clear control flow
Use else for code that should run only when the protected operation succeeded. It keeps that code outside the try, preventing the handler from accidentally catching an unrelated exception:
def load_port(raw: str) -> int:
try:
port = int(raw)
except ValueError as exc:
raise ConfigurationError("port must be an integer") from exc
else:
if not 1 <= port <= 65_535:
raise ConfigurationError("port is outside the valid range")
return port
Use finally for cleanup that must occur whether the operation succeeds, fails, returns, or is cancelled. Do not return from finally: it can replace a normal return value or suppress a pending exception. Python 3.14 emits a SyntaxWarning for return, break, or continue that exits a finally block because the behavior is confusing.
Prefer a context manager when a resource already provides one:
def read_banner(path: Path) -> str:
with path.open(encoding="utf-8") as handle:
return handle.readline().rstrip("\n")
The file closes when the block exits, including when reading raises. A manual finally remains appropriate for resources without a suitable context manager or for cleanup spanning several operations. The official finally and predefined cleanup documentation explains both behaviors.
Log an exception once at the owning boundary
logger.exception() emits an error-level log record and attaches the active exception information. The Python logging reference says it should only be called from an exception handler:
import logging
logger = logging.getLogger(__name__)
class RecordImportError(RuntimeError):
pass
def run_import(job_id: str) -> int:
try:
import_records(job_id)
except RecordImportError:
logger.exception(
"record import failed",
extra={
"event": "import.failed",
"job_id": job_id,
},
)
return 1
return 0
This boundary owns the command result, so logging and returning a failure status is coherent. If the handler logs and then re-raises, a caller may log the same traceback again. Prefer one owner: lower layers translate or add structured safe context; the boundary that turns the exception into a failed request, job, command, or message logs it.
logger.error("...", exc_info=True) is the explicit alternative when another error-level method reads more clearly. Inside a handler, it attaches the current exception. The logging API also accepts an exception instance or exception tuple through exc_info, but logging.exception() is harder to misuse for the common case.
Never return the traceback to an HTTP client. Send a generic error code and an opaque correlation identifier that the client is allowed to know. Keep the detailed traceback in a restricted log sink. Stack frames expose source paths and call structure, while the exception message may expose data from the failed operation.
Choose severity from the outcome
An exception type does not determine the log level by itself. A handled parse failure caused by invalid optional configuration might be a warning. A timeout that makes one required job fail may be an error. A validation error returned normally to a client may need an informational audit event rather than a traceback. Reserve critical severity for a condition with the operational meaning your service assigns to it; do not mark every caught exception critical.
Avoid logging the same expected failure twice at different levels. For example, a library should not emit an error and then raise ValueError if its caller is expected to translate that exception into a normal client response. Libraries can expose typed exceptions; the application boundary decides whether the final outcome is expected, degraded, or failed.
Use exception type as a bounded diagnostic field only when it helps grouping. A class name such as TimeoutError is safer than the exception's string, but it is still not a stable product contract: dependencies can change their exception hierarchy between versions. Pair it with a stable application event such as inventory.reserve_failed and document the event's meaning independently of the Python class.
Sampling also needs an outcome rule. Repeated expected validation events may be sampled or counted outside the error stream, while rare unexpected failures should retain enough evidence for diagnosis. Never sample by message text that may contain secrets, and do not sample away required audit evidence. Record dropped or sampled counts separately so a quiet log stream is not mistaken for a healthy service.
Preserve safe structured context with the logging API
Use a configured Formatter, handler, and LoggerAdapter rather than building JSON strings at every catch site. This Python 3.14 example puts stable, allowlisted fields on the LogRecord and writes to standard error through StreamHandler:
import logging
def build_logger() -> logging.LoggerAdapter:
handler = logging.StreamHandler()
handler.setFormatter(
logging.Formatter(
"timestamp=%(asctime)s level=%(levelname)s "
"service=%(service)s environment=%(environment)s "
"event=%(event)s operation_id=%(operation_id)s "
"message=%(message)s",
defaults={"event": "unset", "operation_id": "none"},
)
)
base = logging.getLogger("checkout.worker")
base.handlers.clear()
base.addHandler(handler)
base.setLevel(logging.INFO)
base.propagate = False
return logging.LoggerAdapter(
base,
{"service": "checkout-worker", "environment": "production"},
merge_extra=True,
)
logger = build_logger()
def run_operation(operation_id: str) -> bool:
try:
reserve_inventory()
except TimeoutError:
logger.exception(
"inventory reservation failed",
extra={
"event": "inventory.reserve_failed",
"operation_id": operation_id,
},
)
return False
return True
merge_extra=True is available in Python 3.13 and later; it combines the adapter's stable service fields with safe call-specific fields. Configure handlers once in the application entry point rather than clearing or adding handlers inside an imported library. The example keeps setup beside the function only to show the complete field path.
The traceback appended by the standard formatter is multiline. A collector must assemble that exception as one bounded logical event or preserve it in a dedicated field; otherwise one failure can become many unrelated records. Set a maximum event size and timeout in the collector, and verify the result with a synthetic exception.
Allowlist context rather than serializing an exception or request object. Useful fields can include a stable event name, service, environment, release, operation type, and opaque operation ID. Exclude authorization headers, cookies, session identifiers, request or response bodies, customer attributes, exception .args, and local variables. Hashing a sensitive value does not automatically make it safe or anonymous.
Format tracebacks only when another API needs the lines
The logging module already formats active exception information for normal exception logs. Use traceback when a different API requires a rendered representation or when a test must inspect formatting:
import traceback
def render_exception(exc: BaseException) -> str:
lines = traceback.format_exception(exc, chain=True)
return "".join(lines)
Python 3.14's traceback.format_exception() documentation states that it returns a list of strings and includes chained exceptions when chain=True. Joining the lines recreates the printable traceback text.
Rendered text remains sensitive and may be large. Do not attach it to an HTTP response, metric label, alert title, or unrestricted analytics event. Bound its size at the storage boundary and preserve the structured exception log whenever possible.
Avoid TracebackException(..., capture_locals=True) in production logging. Python's traceback documentation notes that captured locals are shown in the output. Locals can hold credentials, tokens, full payloads, ORM objects, and personal data. The convenience is rarely worth the disclosure risk.
Preserve cancellation and grouped async failures
Python 3.11 introduced asyncio.TaskGroup for structured concurrency. On the first child failure other than CancelledError, the group cancels remaining children. After all tasks finish, it raises the failures in an ExceptionGroup or BaseExceptionGroup as appropriate. KeyboardInterrupt and SystemExit receive special handling. These details are defined in the current TaskGroup documentation.
Use except* when different failure types inside a group need different handling:
import asyncio
class InvalidRecordError(ValueError):
pass
async def import_batch(records: list[str]) -> None:
try:
async with asyncio.TaskGroup() as group:
for record in records:
group.create_task(import_one(record))
except* InvalidRecordError as group:
quarantine_invalid_records(group.exceptions)
An except* clause receives the matching subgroup. Any unhandled exceptions continue through other except* clauses and are re-raised after processing. Handling the subgroup is appropriate only if quarantine is the intended outcome; do not treat a group of partially completed operations as a full success accidentally.
Cancellation is control flow, not an ordinary application failure. asyncio.CancelledError directly subclasses BaseException. If code catches it explicitly to clean up, re-raise it when cleanup completes:
async def consume_messages() -> None:
try:
await consume_forever()
except asyncio.CancelledError:
await close_consumer()
raise
The asyncio cancellation guidance warns that swallowing cancellation can break TaskGroup and timeout behavior because structured-concurrency components use cancellation internally. A finally block is often enough when cleanup does not require special cancellation handling.
Own standalone background tasks
Prefer TaskGroup when tasks share a lifetime. If a task truly outlives the current scope, keep a strong reference and retrieve its result:
import asyncio
background_tasks: set[asyncio.Task[None]] = set()
def report_result(task: asyncio.Task[None]) -> None:
background_tasks.discard(task)
try:
task.result()
except asyncio.CancelledError:
return
except Exception:
logger.exception(
"background task failed",
extra={"event": "background_task.failed"},
)
def start_background_work() -> None:
task = asyncio.create_task(refresh_cache())
background_tasks.add(task)
task.add_done_callback(report_result)
The event loop keeps only weak references to tasks, so the official create_task() guidance recommends retaining a reference for reliable background work. Calling result() after completion returns the result or re-raises the stored exception, which prevents failures from being silently abandoned. The callback logs once, at the boundary that owns the detached task.
Understand thread and top-level exception hooks
threading.excepthook runs for an exception that escapes Thread.run(). By default, it silently ignores SystemExit and prints other uncaught thread exceptions to sys.stderr. A custom hook can route the exception through application logging:
import threading
def log_uncaught_thread_exception(args: threading.ExceptHookArgs) -> None:
if args.exc_type is SystemExit:
return
logger.error(
"thread terminated with an uncaught exception",
extra={"event": "thread.uncaught_exception"},
exc_info=(args.exc_type, args.exc_value, args.exc_traceback),
)
threading.excepthook = log_uncaught_thread_exception
Install the hook once during application startup. Do not save args, args.exc_value, or args.thread for later. The Python threading reference warns that retaining the exception value can create a reference cycle and retaining the thread can resurrect an object during finalization.
sys.excepthook has a narrower role than its name suggests. The interpreter calls it for an uncaught non-SystemExit exception at the top level, just before a program exits or an interactive prompt resumes. It does not replace threading.excepthook, retrieve failed asyncio task results, catch framework-managed request errors, or provide service-wide exception capture. The current sys.excepthook documentation states this top-level boundary explicitly.
Global hooks are final reporting boundaries, not recovery mechanisms. In a non-interactive program, sys.excepthook runs just before exit; in an interactive session, it runs before control returns to the prompt. Keep the hook simple, avoid network-dependent cleanup, and preserve the original failure if custom reporting fails.
Keep client responses separate from internal evidence
At an HTTP or RPC boundary, translate an internal failure into the protocol's safe response without copying the exception message or traceback. A response can include a stable public error code and an opaque operation identifier. The internal log can use the same identifier to correlate the sanitized response with restricted evidence.
Do not assume that every exception is a server error. Expected validation failures can map to a documented client response without an error-level traceback. Dependency timeouts, invariant violations, and unexpected exceptions need different protocol status and retry semantics. Centralize that mapping in the boundary instead of scattering it across business functions.
Logging and re-raising at every layer creates duplicate events and contradictory severity. Translate locally, add only safe typed context, and log once where the program decides the work unit failed. If an infrastructure framework already logs uncaught request exceptions, either configure that handler with the required context or prevent a second application log from duplicating it.
Test the exception path and the log path
Test control flow separately from logging output:
- the expected exception type is caught and the intended fallback or translation occurs;
- an unexpected exception still propagates;
raise ... from excsets the expected cause;- cleanup runs on success, failure, return, and cancellation;
TaskGroupfailures cancel siblings and the expected subgroup is handled;CancelledErroris re-raised after cleanup;- a detached task is retained until completion and its result is retrieved;
- the thread hook logs an uncaught thread exception without retaining hook arguments;
- client responses contain no traceback, secret, or raw exception message.
For logs, attach a temporary in-memory handler in the test rather than asserting terminal text. Inspect the LogRecord: level, event name, service, environment, operation identifier, and whether exc_info is present. Assert that prohibited values—tokens, payloads, personal identifiers, and exception arguments—are absent from the formatted output.
Then test the collection path. Emit a unique synthetic exception in a non-production environment, confirm the collector assembles it as one bounded event, and verify that the destination preserves the safe fields and cause chain. Also test collector rejection, buffering limits, oversize behavior, and a destination outage. Application tests cannot prove that a production logging pipeline delivers every event.
Investigate captured Python exceptions in Fluxtail
Fluxtail is a paid, self-service Starter/Pro logs-focused service. A Python application can write configured logging output to standard output or another supported handler; a collector can then forward those records through a documented Fluxtail receiver into a named stream. The exact parsing and field names depend on the formatter, collector, receiver, and verified mapping.
After one synthetic marker arrives, use search and filters and Live Tail to confirm the event name, service, environment, operation ID, message, and traceback representation that actually landed. Fluxtail does not instrument Python, catch exceptions inside the process, or provide tracing, APM, or metrics. It also cannot recover records dropped before they reach its receiver.
Built-in AI chat and the hosted MCP server are separate investigation interfaces. Hosted MCP uses account-bound OAuth with PKCE; its access remains limited to the authorized account, and mutations are proposed before a short-lived confirmation is applied. Use bounded time windows and verify summaries against retained raw rows. Neither interface replaces safe exception handling or proves the collector delivered every failure.
The broader log management best-practices guide covers source allowlists, buffering, access, retention, and pipeline validation. If a verified Python logging path needs centralized search and retention, create a Fluxtail account and choose the paid plan that fits the ingest volume and retention requirement.
Apply a final exception-capture checklist
Before shipping a handler, confirm:
- The
tryblock contains only the operation whose expected failure is handled. - The handler catches a documented exception type, not bare
exceptorBaseException. - The code recovers, translates, adds safe context, or terminates one owned work unit.
- Unexpected exceptions still propagate to an accountable boundary.
- Cause chains survive translation through
raise ... from exc. - Exactly one boundary logs the failure unless separate events have distinct operational meaning.
logging.exception()runs only inside an active handler.- Context fields are stable, bounded, and allowlisted; exception arguments and request data are excluded.
- Tracebacks stay in a protected sink and never reach client responses.
- Async cancellation is propagated, task lifetimes are owned, and task results are retrieved.
- Thread and top-level hooks are treated as limited reporting boundaries.
- Tests prove both program behavior and the end-to-end collection result.
Good exception capture is intentionally narrow. It preserves control flow, evidence, and causal context while keeping sensitive data out of responses and broad telemetry. That makes failures visible without turning every except block into a second, inconsistent logging system.