Under which category would you file this issue?
Providers
Apache Airflow version
3.2.2
What happened and how to reproduce it?
GitDagBundle authenticated as a GitHub App fails intermittently with remote: Repository not found. immediately after the hook mints a new installation access token. The same fetch against the same URL succeeds unchanged about 10 s later. A newly created installation token is evidently not usable at GitHub's git HTTPS endpoint for a few seconds after POST /app/installations/{id}/access_tokens returns it.
Two code paths are affected:
-
dag-processor refresh(). _fetch_bare_repo() has no retry, so the manager logs Error refreshing bundle and picks the bundle up on its next cycle. Harmless, but it happens on every affected token refresh. Over 7 days across two deployments every one of 23 occurrences was within 1 s of Successfully obtained GitHub App installation access token; none happened mid-lifetime.
13:07:14.005 Refreshing bundle dags
13:07:14.006 GitHub App token is missing or near expiry (expires at: 2026-09-28 13:11:15+00:00). Refreshing token.
13:07:14.659 Successfully obtained GitHub App installation access token (expires at: 2026-09-28 14:07:14+00:00)
13:07:15.406 Error refreshing bundle dags
git.exc.GitCommandError: Cmd('git') failed due to: exit code(128)
cmdline: git fetch -v -- origin +refs/heads/*:refs/heads/* +refs/tags/*:refs/tags/*
stderr: 'fatal: repository 'https://github.com/<org>/<repo>.git/' not found'
13:07:19.224 Refreshing bundle dags
13:07:20.165 Error refreshing bundle dags (same error)
13:07:24.917 Refreshing bundle dags (succeeds)
-
Task startup with KubernetesExecutor. Every task pod mints its own token in GitHook and immediately calls _clone_bare_repo_if_required(). That method is decorated with stop_after_attempt(2) and no wait, so both attempts run inside the bad window (0.9 s apart) and the task dies before it runs:
10:50:15.312 GitHub App token is missing or near expiry (expires at: None). Refreshing token.
10:50:16.104 Successfully obtained GitHub App installation access token (expires at: 2026-09-28 11:50:16+00:00)
10:50:16.106 Cloning bare repository
10:50:16.963 Bare repository clone/open/fetch failed, cleaning up and retrying
10:50:17.695 Bare repository clone/open/fetch failed, cleaning up and retrying
10:50:17.696 Top level error
RuntimeError: Error cloning repository
File ".../airflow/sdk/execution_time/task_runner.py", line 775, in parse
File ".../airflow/providers/git/bundles/git.py", line 247, in initialize
File ".../airflow/providers/git/bundles/git.py", line 203, in _initialize
GitCommandError: Cmd('git') failed due to: exit code(128)
cmdline: git clone -v --bare -- https://github.com/<org>/<repo>.git /tmp/airflow/dag_bundles/dags/bare
stderr: 'Cloning into bare repository ...
remote: Repository not found.
fatal: repository 'https://github.com/<org>/<repo>.git/' not found'
In one production deployment 3 of 51 task-pod clones over 7 days failed this way, each one failing the task instance (and with retries=0, the DAG run).
Reproduce: configure a GitDagBundle with a git connection using github_app_id / github_installation_id / key_file against a private repository, run tasks with KubernetesExecutor (or anything that creates a fresh GitHook per task), and watch task logs. Roughly 1 in 25 fresh tokens hits the window, so it shows up within a day of normal traffic. The repository, the App installation and the network are fine: the same token works seconds later.
Bundle config used:
[dag_processor]
dag_bundle_config_list = [{"name": "dags", "classpath": "airflow.providers.git.bundles.git.GitDagBundle",
"kwargs": {"git_conn_id": "github_dags", "subdir": "dags", "refresh_interval": 60, "tracking_ref": "main"}}]
AIRFLOW_CONN_GITHUB_DAGS='{"conn_type":"git","host":"https://github.com/<org>/<repo>.git",
"extra":{"github_app_id":"...","github_installation_id":"...","key_file":"/etc/git-app/key.pem"}}'
What you think should happen instead?
The bundle should tolerate GitHub's token propagation delay. _clone_bare_repo_if_required (and ideally _fetch_bare_repo when called from refresh()) should retry with a wait, for example stop_after_attempt(5) plus wait_exponential(multiplier=2, max=15), so at least one attempt lands after the token is usable. Alternatively the hook could verify a freshly minted token (for example with git ls-remote) before handing it to the clone. Today the only workaround is task-level retries, which re-runs the whole task in a new pod with a new token.
We are running exactly this change as a build-time patch of 0.5.0 and it removes the startup failures. main still has stop_after_attempt(2) with no wait on _clone_bare_repo_if_required, so the issue is present in 1.0.0rc1 as well.
Operating System
Debian 12 (apache/airflow:3.2.2-python3.10 image)
Deployment
Official Apache Airflow Helm Chart
Apache Airflow Provider(s)
git
Versions of Apache Airflow Providers
apache-airflow-providers-git==0.5.0 (with github extra)
apache-airflow-providers-cncf-kubernetes==10.22.0
Official Helm Chart version
1.16.0
Kubernetes Version
AKS
Helm Chart configuration
dagProcessor.dagBundleConfigList as above, dags.gitSync.enabled: false, KubernetesExecutor, GitHub App private key mounted from a Secret at the connection's key_file path.
Docker Image customizations
Two build-time patches of apache-airflow-providers-git 0.5.0: the ETXTBSY askpass fix from #73425, and the clone retry with backoff described above. The failures in this report were observed before the retry patch was applied.
Anything else?
The propagation delay is on GitHub's side, but since the provider is the one minting the token and using it a few hundred milliseconds later, it is the natural place to absorb it.
Are you willing to submit PR?
Code of Conduct
Under which category would you file this issue?
Providers
Apache Airflow version
3.2.2
What happened and how to reproduce it?
GitDagBundleauthenticated as a GitHub App fails intermittently withremote: Repository not found.immediately after the hook mints a new installation access token. The same fetch against the same URL succeeds unchanged about 10 s later. A newly created installation token is evidently not usable at GitHub's git HTTPS endpoint for a few seconds afterPOST /app/installations/{id}/access_tokensreturns it.Two code paths are affected:
dag-processor
refresh()._fetch_bare_repo()has no retry, so the manager logsError refreshing bundleand picks the bundle up on its next cycle. Harmless, but it happens on every affected token refresh. Over 7 days across two deployments every one of 23 occurrences was within 1 s ofSuccessfully obtained GitHub App installation access token; none happened mid-lifetime.Task startup with KubernetesExecutor. Every task pod mints its own token in
GitHookand immediately calls_clone_bare_repo_if_required(). That method is decorated withstop_after_attempt(2)and nowait, so both attempts run inside the bad window (0.9 s apart) and the task dies before it runs:In one production deployment 3 of 51 task-pod clones over 7 days failed this way, each one failing the task instance (and with
retries=0, the DAG run).Reproduce: configure a
GitDagBundlewith agitconnection usinggithub_app_id/github_installation_id/key_fileagainst a private repository, run tasks with KubernetesExecutor (or anything that creates a freshGitHookper task), and watch task logs. Roughly 1 in 25 fresh tokens hits the window, so it shows up within a day of normal traffic. The repository, the App installation and the network are fine: the same token works seconds later.Bundle config used:
What you think should happen instead?
The bundle should tolerate GitHub's token propagation delay.
_clone_bare_repo_if_required(and ideally_fetch_bare_repowhen called fromrefresh()) should retry with a wait, for examplestop_after_attempt(5)pluswait_exponential(multiplier=2, max=15), so at least one attempt lands after the token is usable. Alternatively the hook could verify a freshly minted token (for example withgit ls-remote) before handing it to the clone. Today the only workaround is task-levelretries, which re-runs the whole task in a new pod with a new token.We are running exactly this change as a build-time patch of 0.5.0 and it removes the startup failures.
mainstill hasstop_after_attempt(2)with nowaiton_clone_bare_repo_if_required, so the issue is present in 1.0.0rc1 as well.Operating System
Debian 12 (apache/airflow:3.2.2-python3.10 image)
Deployment
Official Apache Airflow Helm Chart
Apache Airflow Provider(s)
git
Versions of Apache Airflow Providers
apache-airflow-providers-git==0.5.0 (with
githubextra)apache-airflow-providers-cncf-kubernetes==10.22.0
Official Helm Chart version
1.16.0
Kubernetes Version
AKS
Helm Chart configuration
dagProcessor.dagBundleConfigListas above,dags.gitSync.enabled: false, KubernetesExecutor, GitHub App private key mounted from a Secret at the connection'skey_filepath.Docker Image customizations
Two build-time patches of
apache-airflow-providers-git0.5.0: the ETXTBSY askpass fix from #73425, and the clone retry with backoff described above. The failures in this report were observed before the retry patch was applied.Anything else?
The propagation delay is on GitHub's side, but since the provider is the one minting the token and using it a few hundred milliseconds later, it is the natural place to absorb it.
Are you willing to submit PR?
Code of Conduct