* AttributeError when start_event is None, and guard process faster stop
This commit has two related fixes that come into play when frequently
shutting down and restarting the cluster.
If a SIGINT or SIGTERM signal is received while the main process is
waiting for the sentinel process to set start_event, the signal
handler sets start_event to None, and the loop that polls start_event
will raise an unhandled AttributeError attempting to check
start_event.is_set().
The second issue is that the guard process sleeps for the guard cycle
interval in each iteration before checking the stop event, preventing
the guard cycle from terminating until the sleep completes. The
cluster can be made more responsive to a shutdown request by using
the event's wait(cycle) method rather than time.sleep(cycle) so the
process wakes immediately when the event is set.
The commit also fixes the case where the guard process sleeps for an
extra cycle when the cycle counter happens to have been reset to zero
on the loop iterations when the scheduler is called.
This fix helps in our environment where we are frequently stopping
and starting the cluster and where we have increased the guard
cycle setting to several seconds.
* Add sentinel premature death check while waiting for sentinel to start.
Suggested by github copilot [here](https://github.com/django-q2/django-q2/pull/305/changes#r3143964420)
* Add test case test_cluster_early_stop
Add a test case to validate stopping the cluster before the
sentinel has set the cluster's start_event.
The test deterministically triggers the AttributeError when
start_event is None (without the PR's fix in cluster.py).
As requested by copilot: https://github.com/django-q2/django-q2/pull/305#discussion_r3143964399
* Add test test_cluster_stop_responsive
Ensure that stopping the cluster is responsive and does not wait for a full
GUARD_CYCLE to stop the cluster.
As requested by copilot: https://github.com/django-q2/django-q2/pull/305/changes#r3143964414
* Address copilot review comments.
1. Do not raise RuntimeError if sentinel early exit was caused by
SIGINT/SIGTERM stopping the cluster while waiting for the
sentinel to start.
Addresses https://github.com/django-q2/django-q2/pull/305#discussion_r3298618046
2. Ensure that the monkeypatched broker instance remains pickleable
so the test_cluster_early_stop test will also work on platforms
that use the spawn Process start method.
Addresses https://github.com/django-q2/django-q2/pull/305#discussion_r3298618086
3. Do not allow the test_cluster_early_stop test to hang if the
sentinel process or main test process terminates unexpectedly
without setting the sentinel_event or test_event.
Addresses https://github.com/django-q2/django-q2/pull/305#discussion_r3298618102
* feat: option to save tasks per group/func/name
* test: update tests for save limit
* fix: convert func when save-limit checking
when passed as function it needs to be converted to work
* Add warning if option is not valid
Co-authored-by: Noortheen Raja <jnoortheen@gmail.com>
* Black linting
* Adding Black to dev dependencies and upping minimal python to 3.6.2 for compatibility
* Updating packages
* Removing pip-tools input and exporting requirements with poetry
* Deleting old setup files and test runner
* Trying 1.3.7
* Looser extras requirements to prevent conflicts
* Sorted imports with isort
* Added iSort to dev dependencies
* Fixes localtime for naive setups
* Some small refactors
* Initial test script
* Fixes django matrix
* Trying to set up Disque
* Disque as a docker service
* lowercase action name
* Remove Travis
* Replaces Travis badge with Github actions badge
* Show django version in step
* Typo
The cluster timeout configuration default value None is documented to mean
tasks never timeout out. Also the documentation states that the timeout
can be overridden for individual tasks.
With the old implementation the timeout parameter given to async_task was
not honored if the cluster timeout was set to None.
Fixes: #335
The old implementation used count_forever test task that never finishes.
If the timeout implementation does not work in the test, the test never
end and you was required to kill the test runner.
* async is now a reserved word in python3.7
* Rename async function to enqueue
* Rename all async_ functions to enqueue_
* Rename Async class to AsyncTask
* Updates the docs.
If a task fails with an exception, it is retried until it
succeeds. This is contrary to what is said in the documentation: under
the "Architecture" section, heading "Broker" it says that even when a
task errors, it's still considered a successful delivery. Failed tasks
never get acknowledged however, thereby being retried after the
timeout period. See also issues #238 and #194.
This patch adds an option to acknowledge failures, thereby closing
issue #238. Issue #194 would require some more work. The default of
this option is set to `False`, thereby maintaining backwards
compatibility.
The test suite, especially for the brokers, was heavily dependant on the
success of the tests preceeding it to pass. This commit should eliminate
that dependency and allow tests to fail without affecting others.
It also removes the need to cleanup globals manually after a test.