Repository navigation
feat(celery): allow custom monitor config for beat tasks - #7839
alexon1234 wants to merge 1 commit into
Conversation
Celery Beat auto-instrumentation derives the cron monitor config from the schedule only, so max_runtime, checkin_margin, failure_issue_threshold, recovery_threshold and owner cannot be set. Tasks whose runtime can exceed Sentry's default 30-minute max_runtime are then reported as timed-out cron failures even when they complete successfully. Add a `beat_task_monitor_config` option to CeleryIntegration that maps beat task names to partial monitor configs, merged over the derived config. Refs getsentry#7838
ericapisani
left a comment
There was a problem hiding this comment.
Thanks for your contribution!
This is a great start, I've left some notes that we'll need to tackle before this gets merged in. Let me know if you have any questions about what I've left here.
| propagate_traces: bool = True, | ||
| monitor_beat_tasks: bool = False, | ||
| exclude_beat_tasks: "Optional[List[str]]" = None, | ||
| beat_task_monitor_config: "Optional[Dict[str, MonitorConfig]]" = None, |
There was a problem hiding this comment.
Since we need a monitor_name provided as part of the configuration for any overrides to work, we can set the default value to an empty dictionary instead of None here.
| beat_task_monitor_config: "Optional[Dict[str, MonitorConfig]]" = None, | |
| beat_task_monitor_config: "Optional[Dict[str, MonitorConfig]]" = {}, |
This would allow us to update the (integration.beat_task_monitor_config or {}).get(monitor_name) conditional to either:
integration.beat_task_monitor_config.get(monitor_name)or
getattr(integration.beat_task_monitor_config, monitor_name, None)depending on your preference.
| assert monitor_config["schedule"] == override_schedule | ||
|
|
||
|
|
||
| def test_beat_task_monitor_config_option(): |
| fake_scheduler = MagicMock() | ||
| fake_scheduler.apply_entry = fake_apply_entry |
There was a problem hiding this comment.
Within the SDK, we're trying to move towards testing functionality via the public APIs rather than the private methods and use mocks as sparingly as we can.
Instead of this approach, can we look to instead do something similar to the test_beat_task_crons_success in test_celery_beat_cron_monitoring.py (which creates a celery app and asserts that the envelopes contain the right data)?
| } | ||
|
|
||
|
|
||
| def test_get_monitor_config_overrides_can_replace_schedule(): |
There was a problem hiding this comment.
We shouldn't support overriding the schedule (or its type) via this new property as there's validation that we do within the SDK to ensure the provided values are valid.
Instead, let's do the following:
1️⃣ Introduce a new type in _types.py to a subset of what's needed to support long-running tasks, etc. and update the relevant spots in this PR
MonitorConfigOverrides = TypedDict(
"MonitorConfigOverrides",
{
"checkin_margin": int,
"max_runtime": int,
"failure_issue_threshold": int,
"recovery_threshold": int,
"owner": str,
},
total=False,
)
2️⃣ Validate the overrides in the __init__ of the integration to ensure that everything's valid
In celery/beat.py:
_ALLOWED_MONITOR_CONFIG_OVERRIDE_KEYS = frozenset(
{
"checkin_margin",
"max_runtime",
"failure_issue_threshold",
"recovery_threshold",
"owner",
}
)
def _validate_beat_task_monitor_config(
beat_task_monitor_config: "Dict[str, MonitorConfigOverrides]",
) -> None:
for task_name, overrides in beat_task_monitor_config.items():
invalid = set(overrides) - _ALLOWED_MONITOR_CONFIG_OVERRIDE_KEYS
if invalid:
raise ValueError(
f"Unsupported keys in beat_task_monitor_config for '{task_name}': "
f"{sorted(invalid)}. Allowed keys: "
f"{sorted(_ALLOWED_MONITOR_CONFIG_OVERRIDE_KEYS)}"
)and then invoke _validate_beat_task_monitor_config(beat_task_monitor_config) in the init code in __init__.py before assigning it to self.beat_task_monitor_config on CeleryIntegration.
Description
Celery Beat auto-instrumentation (
CeleryIntegration(monitor_beat_tasks=True)) derives the cron monitor config from the schedule only (schedule+timezone), so there is no way to setmax_runtime,checkin_margin,failure_issue_threshold,recovery_threshold, orowneron the monitors the SDK creates.Consequence: a task that can legitimately run longer than Sentry's default 30-minute
max_runtimeis reported as a timed-out cron failure even though it completes successfully. The only workaround is to exclude it from auto-instrumentation and re-declare the schedule in a manual@sentry_sdk.monitor(..., monitor_config=...)call, which is awkward when the schedule differs per environment.This adds a
beat_task_monitor_configoption toCeleryIntegration: a mapping of Celery Beat task name to a partialMonitorConfigthat is merged over the config derived from the schedule.The SDK still derives
scheduleandtimezone; the override only needs the extra fields, and can still override the derived values if explicitly provided.Changes:
CeleryIntegration: newbeat_task_monitor_configoption._get_monitor_config: accepts and merges monitor config overrides._apply_crons_data_to_schedule_entry: looks up the override for the schedule entry and passes it through (covers both the default scheduler and RedBeat, which share this path).Issues
Reminders
uv run ruff.feat:,fix:,ref:,meta:)