Trello API Outage History
Past incidents and downtime events
Complete history of Trello API outages, incidents, and service disruptions. Showing 50 most recent incidents.
May 2026(4 incidents)
Trello - degraded performance
19 updates
# Summary On May 17th 2026 at 12:00 UTC a subset of Trello users experienced a degradation of some core functionality including but not limited to: issues with comments, viewing activity logs, accessing card details, board exports, Power-Ups, attachments, Home page updates, and search functionality. Some moved cards displayed 'Card Not Found' errors. These issues persisted for multiple days. Most of the issues encountered when interacting with Trello (commenting, viewing activity) were resolved by May 19 2026 at 08:31 UTC, however there was a lingering data integrity issue which manifested as “Card Not Found” errors when trying to open cards. >90% of these cards were repaired by May 28 at 17:20 UTC. The remainder took a little more time and were completely repaired by June 5th 2026 at 21:37 UTC. We wanted to share a more detailed analysis of the issues we encountered, as well as what we’re doing to ensure nothing similar happens in the future. # Root Cause The incident was caused by a bug in our database (MongoDB) which surfaced during a project to optimize the performance and scale of one of our largest collections. The affected collection contains card comments and other activity and is used across much of Trello’s functionality. The optimization project was to reshard the collection onto a new key. This allows us to eliminate a so-called “insert hot spot” which is where all newly created database entries end up on a single database partition (aka shard). After resharding, the load would be spread evenly across the various servers that make up the database cluster. This decreases latency and improves query performance for users. During the final stages of the resharding operation (which is a long-running, automated process), 12 of the 30 relevant shards ran out of disk space while building indexes. Rather than aborting the resharding process, the database bug caused the cluster to put the new database servers into rotation without their indexes, despite logs and metrics saying otherwise. Due to the enormous amount of data in this collection, having no indexes means that there is essentially no way to query for the data you want (i.e. “give me all comments for this card”) without performing an extremely slow and inefficient collection scan. # Remedial Actions Plan The first thing we did was initiate a full rebuild of all the database indexes for the collection. Unfortunately this takes a very long time to complete (multiple days) and is compute-intensive. We scaled all the database servers up in order to help the process move faster, however due to the single-threaded nature of this process, more compute did not equate to more speed. Additionally we enabled a database setting which caused all queries that would result in a full collection scan to fail-fast rather than drag on in the background and consume database resources before ultimately failing. This helped the parts of Trello that were still working stay responsive. Once the reindex process had completed, all functionality was restored. However, we discovered that cards which were moved to a different board during the rebuild were left in a broken state. Their activity logs and comments were essentially stranded on the old board while the card was now on the new board. This manifested as a “Card Not Found” error when trying to open such a card. The team moved into a new phase of writing code to safely remediate all of the remaining data issues in batches, with the final batch of fixes completing by June 5th 2026 at 21:37 UTC. # Next Steps We know that Trello is where your work lives and issues like this can cause a lot of disruption. To ensure we don’t run into this issue again in the future, the team is taking the following actions: - Working closely with our database vendor to ensure the original bug is fixed (update: the fix is shipping in the next version) - Working to create a standard set of tools and processes to ensure index rebuilds proceed as fast as possible in an emergency like this - Designing and implementing changes to make card move operations more atomic (either the whole operation succeeds, or the whole operation fails, rather than a partial failure)
Thank you for your patience as our team worked diligently to resolve outstanding issues. Although our team encountered edge scenarios through the weekend, we believe all 'card not found' issues are resolved. Minor data inconsistencies listed below may be present if cards were moved between boards during the incident: - Missing card labels - Attachment restrictions not being enforced on moved cards - Checklist assignee information on some cards - Incorrect due date reminders sent to some card assignees Please be assured that our team continues the work necessary to resolve above issues; however, given the impact and scope of remaining issues, we are now moving the channel of communication to Support rather than this Statuspage. If you believe that you continue to see lingering impacts from this incident, please do not hesitate to reach out to our Support team for further assistance.
Our team is still working to fix minor data inconsistencies on cards that were moved between boards during the incident. The following are still being addressed: -Missing card labels -Attachment restrictions are not being enforced on moved cards -Checklist assignee information on some cards. -Incorrect due date reminders sent to some card assignees Your cards are safe to edit. If you make changes to a card before our automated fix reaches it, your edits will be preserved, and the automated restore for that card will be skipped. No data will be lost in either scenario. We will provide another update once the above items are fully resolved. If you need assistance in the meantime, please reach out to our support team.
The team continues to remediate potential outstanding data issues that occurred on Boards and Cards during the incident
Users should be experiencing restored functionality to their boards and cards. Our team has resolved 'card not found' errors, and continues to remediate potential outstanding data issues that occurred on Boards and Cards during the incident: - Missing card labels - Attachment restrictions not being enforced on moved cards - Checklist assignees that are not a member of the current board may incorrectly linger - Card assignees that are not a member of the current board may have incorrectly received due date reminders Please note that while cards are safe to edit, they may be excluded from certain restoration work if they are edited. We will provide another update when above issues are resolved. In the meantime, please reach out to our support team if you need further assistance.
The team continues to remediate potential outstanding data issues that occurred on Boards and Cards during the incident (details are in yesterday's update). Users should be experiencing restored functionality to their boards and cards. Please note that while cards are safe to edit, they may be excluded from certain restoration work if they are edited. We greatly appreciate your patience as we work through these issues. Please reach out to our support team if you need further assistance.
Users should be experiencing restored functionality to their boards and cards. The team continues to remediate potential outstanding data issues that occurred on Boards and Cards during the incident. Below is a list of remaining issues that may still be occurring, but we are currently working to resolve via processing on our end. These issues were caused by card moves between boards during the incident. - Missing card labels - Attachment restrictions not being enforced on moved cards - Checklist assignees that are not a member of the current board may incorrectly linger - Card assignees that are not a member of the current board may have incorrectly received due date reminders Please note that cards are safe to edit, but they may be excluded from certain restoration work if they are edited. We greatly appreciate your patience as we work through these issues. We will provide another update when above issues are resolved. In the meantime, please reach out to our support team if you need further assistance.
Users should be experiencing restored functionality to their boards and cards. We are running scripts to replay actions that occurred during the incident. We will update once these are complete.
Index rebuild is complete and functionality should be restored for users. We are aware that cards moved by users during the incident may be in an incomplete state and we are investigating solutions. We have no indication of data loss during the incident. We will continue to monitor and will provide the next updates in 4 hours or sooner.
We are continuing to monitor the index rebuild process across all affected shards and will provide updates as each phase completes. We will provide next update in eight hours or sooner as we make progress.
We want to provide more transparency regarding this incident given its duration. We were optimizing one of our largest databases, specifically the one storing comments and card activity. At the end of this process, a bug caused some database partitions to go live without their indexes. This means queries to those partitions are timing out, which is why approximately 40 percent of boards are experiencing failures. Your data is safe. Any comments added during this incident are being saved and will reappear once the process is complete. We are currently rebuilding the missing indexes. We have allocated significant compute resources to this task to move as quickly as possible, but the data volume is enormous. We estimate it will take approximately 20 more hours to reach full resolution. Most impacted features include, but are not limited to: - Comments (viewing, adding, notifications) - Activity views on cards - Attachments, Power-Ups, board exports - Home page / 'Updates' view - Search - Some cards showing 'Card Not Found' - Moving/copying cards or lists Please avoid moving or copying cards between boards for now. If your board is working normally, you can continue your work, as most boards remain unaffected. We know Trello is where your work lives and we are committed to getting everything back to normal. We will provide next update in eight hours or sooner as we make progress.
Our teams continue to work diligently on the mitigation. During this time, we recommend that you do not move or copy cards or lists between boards. For users whose boards are affected by this issue, you may create and delete cards and edit card descriptions as usual, but you will not be able to see comments or the detailed activity log of the cards. We estimate that the rebuild will take another 24 hours to complete. We will update this page as we have more information.
Trello continues to experience slowdowns and failures that are affecting comment operations, including creating and deleting comments. Our teams continue to work on the issue resolution, while it is taking longer than expected. We will provide more updates within the next 8 hours or sooner once we have further information to share.
Trello is experiencing slowdowns and failures affecting comment operations, including creating and deleting comments. Users may also experience delays when loading activity views on cards, viewing attachments, using Power-Ups, exporting boards and 'Updates' on home page. Our team is actively working on the resolution and will likely take up to 7 hours to restore the services back to full functionality. We will provide further update at this time, or earlier if we have additional information to share prior to this resolution.
Trello is experiencing slowdowns and failures affecting comment operations, including creating and deleting comments. Users may also experience delays when loading activity views on cards, viewing attachments, using Power-Ups, or exporting boards. Some users are encountering errors when viewing specific cards. The team is working hard to restore full functionality and will provide another update within 3h.
Work by the Trello engineering team continues to restore functionality to comments and activity views. Recovery of the affected systems is taking longer than anticipated. Next update will be posted within 4h or on a significant change. Thank you for your patience.
Work by the Trello engineering team continues to restore functionality to comments and activity views. Next update will be posted within 4h or on a significant change. Thank you for your patience.
The team has identified the root cause and is hard at working resolving this issue. Users will see problems creating/updating/deleting comments and loading the activity view for cards at this time. The rest of Trello's features should be working as normal. We will provide more updates within 2h.
Users may be experiencing slowdowns across various functions of Trello including creating/deleting comments, editing cards, or creating new attachments. The team is actively working to fix the issue and will provide more updates within the hour.
Trello - degraded performance (resolved)
5 updates
Search is now fully up to date for all users, and all issues with previous degraded functionality have been resolved.
The team has resolved the previous issues and is now monitoring the system to ensure a full return to baseline. Search indexing is still approximately 1h delayed, but will be up to date shortly. Thank you for your patience
The Trello team has identified and mitigated most of the issues with cards and comments. Users may still see delays in search indexing for short time while it catches up to realtime.
Trello engineers are actively working on resolving issues that our users may be seeing around commenting on cards, deleting comments, and adding new attachments. Users may also find that recent changes to cards are not reflected in search results. We are continuing to work on it and will provide more updates within the hour.
We are investigating degraded performance of some Trello features. We will provide more details within the next hour.
Users experiencing issues accessing multiple Atlassian products
7 updates
### Summary On May 14, 2026, between 04:30 and 05:26 UTC, Atlassian customers experienced widespread service disruption across multiple Atlassian Cloud products. The issue was caused by a race condition in our internal deployment orchestration platform during a routine rollback operation of a core identity service in the us-east region. This race condition resulted in insufficient capacity for the identity service in the affected region which started returning errors to dependent products. The incident was detected within a minute by automated monitoring systems and mitigated in 56 minutes. ### **IMPACT** During the incident, customers attempting to access Atlassian Cloud products in the us-east region experienced authentication and permission failures and were unable to access services. Customers also experienced errors when accessing the support portal until Atlassian fell back to an alternate support method. This was caused by a core identity service in the us-east region becoming unavailable. Affected products included Atlassian Administration, Atlassian Analytics, Bitbucket, Compass, Confluence, Jira, Jira Product Discovery, Jira Service Management and Trello. Some users outside us-east may have been affected in certain scenarios. ### **ROOT CAUSE** The incident was caused by a race condition in our internal deployment orchestration platform during a routine rollback operation of a core identity service in the us-east region. This race condition resulted in insufficient capacity for the identity service in the affected region which started returning errors to dependent products. ### **REMEDIAL ACTIONS PLAN & NEXT STEPS** We know that outages impact your productivity. Atlassian is prioritizing the following actions to help prevent similar incidents in future: * **Refine deployment orchestration safeguards** * Harden our deployment platform to prevent similar race conditions or resulting capacity loss during a rollback operation. * Streamline mitigation steps when a service becomes unavailable in a region. * **Reduce cross-region impact** * Improve regional isolation and fallback handling so an issue affecting a single region is less likely to impact customers or product functionality in other regions. We recognise how critical reliable access to Atlassian products is for our customers' productivity, and we apologize to customers who were impacted by this incident. Thanks, Atlassian
All products and services impacted by this incident should now be fully recovered, and this incident is resolved.
We are approaching full system recovery at this time, and are performing final confirmations that services are restored.
We are now able to see recovery for all impacted products, and users should be able to access their products as expected. Our team is continuing to monitor all products and services to ensure there is no further impact, and we will provide further update when this has been validated and the incident is closed.
Our team has implemented a mitigation for this issue and we are now seeing recovery across Atlassian products. We will continue to monitor this issue for any ongoing concerns, and provide further updates here within an hour as we are able to confirm a full recovery has taken place.
Our team has identified the root cause of this issue and is now actively working on mitigating the issue with accessing Atlassian products. At this time, Atlassian customers should also be able to once again raise support requests with our team. We will provide further update within an hour as we are able to progress mitigating this issue.
It is likely if you are experiencing any issues relating to logging in or accessing Atlassian products at this time it is likely due to this ongoing incident. We are continuing to receive reports about expanded product impact resulting from this incident. While our team continues to investigate the issue with urgency, we will continue to provide further updates here with additional information. We will provide further update within one hour, or sooner as further information becomes available.
Multiple Atlassian services are experiencing issues
12 updates
All dates and times below are in UTC unless stated otherwise. ### Summary On May 8, 2026 between 00:22 and 06:08, one of our hosting providers suffered a significant incident in a specific availability zone in prod-east which led to Atlassian customers experiencing degraded performance and delays of background operations and automation execution. The incident started on May 8, 2026 at 00:22 and was detected within 4 minutes by automated monitoring systems. Our teams worked to restore core access by 06:08. Final cleanup of backlogged processes and minor issues progressed in stages from there was completed iteratively by 19:15. ### **IMPACT** The primary infrastructure affected in this incident was the event processing pipeline in the prod-east region, which distributes events between Atlassian services and underpins background operations such as automation execution, search indexing, notifications, permission synchronisation. * Between 00:22 and 06:08, an infrastructure incident in our hosting provider triggered an ingestion failure in our event processing pipeline. * At 02:50, event ingestion was failed over to an unaffected availability zone, progressively restoring live event flows. * At 06:08, reliability for new ingestion in prod-east recovered to 100%. The remaining work was to drain the accumulated cross-region backlog of messages, which completed by 17:00. * By 18:48, Automation had processed their backlog of events that were created while processing **Automation** Between 00:22 and 02:50, customers with automation rules triggered by events originating from the prod-east region experienced a significant reduction in rule executions. During this window, event-triggered automation rules were not firing because the events that trigger them were not being delivered. Rule authoring, saving, and rules triggered manually, by schedules, or by webhooks were not affected. At 02:50, the event processing infrastructure failed over to an unaffected availability zone, restoring delivery of live events to Automation and allowing new event-triggered rules to begin executing normally. However, events generated during the impact window still needed to be replayed before delayed automations could be processed. Beginning at 08:28, upstream services replayed their queued events in a coordinated sequence, and all replayed events were processed by 18:48. During the replay window, customers may have experienced automation rules executing later than expected, a small number of rules reaching daily processing limits due to compressed replay, and time-sensitive rules not completing as expected if internal timeout thresholds were exceeded. **Jira and Jira Service Management** Between 00:22 and 02:50, customers with tenants hosted in the prod-east region experienced disruption to Jira and Jira Service Management event-driven features like automation, along with a short period of elevated errors during infrastructure failover. Core Jira experiences, including issue view, boards, and project navigation, remained available throughout the incident. Jira event delivery was affected by the primary impact, preventing downstream services from receiving issue lifecycle events. This affected automation rules triggered by Jira events, AI agent orchestration in Jira, notifications for issue updates and transitions, search indexing for newly created or modified issues, and event-driven integrations between Jira and other Atlassian products. At 02:50, the event processing infrastructure failed over to an unaffected availability zone, restoring delivery of new events. All events generated during the impact window were retained in a recovery queue and required replaying. This began at 08:28 and completed at 12:00. During the replay window, customers may have experienced automation rules executing later than expected, delayed notifications arriving hours after the triggering action, temporary gaps in search results for content created or modified during the impact window, and AI agent workflows not completing as expected where internal timeout thresholds were exceeded. **Confluence** Between 00:22 and 02:50, customers with tenants hosted in the prod-east region experienced disruptions to event-driven services in Confluence. This resulted in delays to search indexing, notifications, automation rule execution, and permission synchronisation. The underlying event processing infrastructure failed over to an unaffected availability zone, after which live Confluence operations resumed normally. However, events generated during the impact window were queued for replay, and some background services remained delayed until that replay and related validation work completed. Between 10:14 and 17:00, a bulk replay of all the queued tenant replay tasks was completed to restore data consistency. During and immediately after the replay window, customers may have experienced search results not reflecting content created or modified during the outage, delayed or missing notifications for page and comment activity, automation rules firing later than expected, and brief delays in permission synchronisation for tenants relying on incremental identity sync. **Bitbucket and Pipelines** Between 00:22 and 06:08, customers using Bitbucket and Pipelines experienced failures and degraded functionality across event-driven workflows. Core Git operations, including push, pull, and clone, were not affected and continued to operate normally throughout the incident. Automatic pipeline triggers initiated by push or pull request events were unavailable during the impact window. Merge queues, custom merge checks, Forge-based triggers, workspace permission changes, and some workspace provisioning flows were also affected. Customers using merge queues were unable to merge pull requests, and some pipeline steps failed because queued work contributed to elevated concurrency limits. At approximately 03:57, Pipelines was reconfigured to consume events through an alternative path, restoring automatic pipeline triggering. Merge queues, custom merge checks, Forge triggers, and other affected workflows were progressively restored as the underlying event processing infrastructure recovered. All Bitbucket and Pipelines services were confirmed fully operational by 06:08. After recovery, queued events were reviewed and replayed where safe to restore data consistency for billing, audit logging, and other background processes. **Identity Services** Between 00:22 and 02:50, customers with tenants hosted in the prod-east region experienced delays in the propagation of identity and group membership changes to downstream Atlassian products. Core identity operations, including authentication, login, and direct group management actions, were not affected and continued to function normally throughout the incident. The impact was limited to asynchronous, event-driven operations that depend on the event processing pipeline. This included delays in delivering group membership and user profile changes to products such as Jira and Confluence, which affected downstream permission synchronisation and crowd sync flows. A small number of SCIM-based identity synchronisation and site provisioning workflows also experienced temporary delays. After the event processing infrastructure recovered, backed-up identity and group directory events were replayed where required, restoring downstream consistency for affected products. No identity data was lost. Group membership changes, user profile updates, and provisioning-related events that occurred during the impact window were retained and processed after recovery. ### **REMEDIAL ACTIONS PLAN & NEXT STEPS** We know outages impact your productivity. While our monitoring and recovery processes helped us respond quickly, this incident highlighted opportunities to further strengthen resilience for event-driven services. We are prioritizing improvements that will: * **Enhance failover coverage** so critical event processing can recover more smoothly during infrastructure disruptions. * **Strengthen recovery handling** so replayed events can be processed more quickly. We apologize to customers whose services were impacted during this incident; we are taking immediate steps to improve the platform’s performance and availability. Thanks, Atlassian Customer Support
This incident has been resolved.
We are continuing to monitor for any further issues.
We continue to monitor the situation as services recover. We are currently in the process of clearing the backlog of queued events. We will provide a further update in approximately one hour.
The underlying issue in public infrastructure which affected asynchronous event processing has been mitigated and all the affected services are recovering. We are now working on clearing the backlog of queued events, which means some actions (such as notifications, automation triggers, and data syncs) may be in degraded state. We will continue to monitor and provide updates as the backlog is cleared.
We are continuing to work with our public cloud provider to mitigate this issue. We are starting to see some recovery in regions outside of Eastern USA, however, users globally may still be experiencing issues with certain product features. These are listed at the bottom of each product page.
Our teams continue to work on mitigating the infrastructure outage from our public cloud provider. We will provide further updates when they are available.
We have identified that the root cause of the issue is related to an infrastructure outage from our public cloud provider. We are working closely with them to mitigate this issue. We will provide further updates when they become available.
Our teams continue to work on mitigating the infrastructure outage from our public cloud provider. We will provide further updates when they are available.
We have identified that the root cause of the issue is related to an infrastructure outage from our public cloud provider. We are working closely with them to mitigate this issue. We will provide further updates when they become available.
We have identified that the root cause of the issue is related to an infrastructure outage from our public cloud provider. We are working closely with them to mitigate this issue. We will provide further updates when they become available.
We are experiencing issues with multiple Atlassian products. Our teams are investigating further and more updates including will be shared within 1 hour.
February 2026(3 incidents)
Trello Slow or down for users
2 updates
Impact Trello users are experiencing issues with websocket connections, leading to difficulties in maintaining active sessions. Affected users may notice disruptions in service as connections are intermittently dropping. Current Status Efforts are underway to restore normal service. Teams are actively engaged in identifying the best path to full resolution. Next Steps The incident team is focusing on pinpointing the root cause and implementing necessary fixes. The next communication will be shared in 30 minutes.
Impact Trello users are experiencing issues with websocket connections, leading to difficulties in maintaining active sessions. Affected users may notice disruptions in service as connections are intermittently dropping. Current Status Efforts are underway to restore normal service. Teams are actively engaged in identifying the best path to full resolution. Next Steps The incident team is focusing on pinpointing the root cause and implementing necessary fixes. The next communication will be shared in 30 minutes.
Degraded performance of Trello
3 updates
On February 19, 2026, Trello users may have experienced performance degradation. The issue has now been resolved, and the service is operating normally for all affected customers.
The issue has been resolved, and services are now operating normally for all affected customers. We'll continue to monitor closely to confirm stability.
We are actively investigating reports of performance degradation affecting Trello. We will share updates here as more information is available.
Trello performance is degraded
3 updates
### **Summary** On Feb 9, 2026, between 07:12 UTC and 08:05 UTC, Atlassian customers were unable to access Trello. The event was triggered by Trello servers reaching maximum memory limits and a subsequent failure to automatically scale for the traffic. The incident was detected within 5 minutes by automated monitoring systems, engaging Trello teams for resolution. Two parallel mitigation efforts were undertaken to restore access: 1\) manual scaling to address concerns with high utilization of hosts, and 2\) temporarily limiting traffic from free customers. This intervention put Trello systems in a known good state with a total time to resolution of ~53 minutes. ### **IMPACT** The overall impact occurred on Feb 9, 2026 between 07:12 AM UTC and 08:05 AM UTC for Trello customers and caused service disruption for all customers, making Trello inaccessible during that time. Access was restored for paid users approximately 43 minutes after the onset, with full service restoration for free users 10 minutes later. ### **ROOT CAUSE** The issue was cause by increased memory requirements during the transition from the weekend to EU business hours, combined with Auto Scaling group scaling rules based on CPU utilization. Memory allocation on our instances outpaced the CPU usage, which caused processes to hit Out-Of-Memory \(OOM\) errors and crash before reaching CPU usage thresholds that would trigger the auto scaling policies. As a result, Trello went down and users received HTTP 502 errors until the incident was resolved. ### **REMEDIAL ACTION PLAN & NEXT STEPS** We know that outages impact your productivity. While we have a suite of automated infrastructure management in place, this specific issue required manual intervention to restore Trello access to users. To avoid repeating this type of incident, we are prioritizing the following remedial action items: * Pre-scale capacity before EU morning traffic - This change has already been introduced to prevent further incidents while we implement additional safeguards. * Adjust Trello’s Auto Scaling group settings - This change will ensure that unhealthy hosts will be replaced more rapidly and scaling policies will consider memory usage. * Refine Trello host memory commitments - This change will decrease the likelihood of memory overcommitment on hosts and associated OOM errors. * Increase isolation for OOM errors - This change will improve the ability for hosts to recover in the event of a single worker experiencing an OOM error. We apologize to customers whose services were impacted during this incident; we are taking immediate steps to improve the platform’s performance and availability. Thanks, Atlassian Customer Support
Trello users experienced performance degradation. The issue has now been resolved, and the service is operating normally for all affected customers.
Trello performance was degraded and the performance degradation of Trello has been resolved. All the services are now operating normally for all affected customers. We'll continue to monitor performance closely to confirm stability.
December 2025(2 incidents)
Outbound Email, Mobile Push Notifications, and Support Ticket Delivery Impacting All Cloud Products
3 updates
### Summary On **December 27, 2025**, between **02:48 UTC and 05:20 UTC**, some Atlassian cloud customers experienced failures in sending and receiving emails and mobile notifications. Core Jira and Confluence functionality remained available. The issue was triggered when **TLS certificates used by Atlassian’s monitoring infrastructure expired**, causing parts of our metrics pipeline to stop accepting traffic. Services responsible for email and mobile notifications had a critical path dependency on monitoring path leading to service disruptions. All impacted services were fully restored by **05:20 UTC**, around **2.5 hours** after customer impact began. ### IMPACT During the impact window, customers experienced: * **Outbound product email failures** \(notifications and other product emails did not send\). * **Identity and account flow failures** where emails were required \(e.g. sign‑ups, password resets, one‑time‑password / step‑up challenges\). * **Jira and Confluence mobile push notifications** * **Customer site activations and some admin policy changes** failing and requiring later reprocessing. ### ROOT CAUSE The incident was caused by: 1. **Expired TLS certificates** on domains used by our monitoring and metrics infrastructure caused by **misconfigured DNS authorization record** which prevented automatic renewal. 2. **Tight coupling of services to metrics publishing**, which caused them to fail when monitoring endpoints became unavailable, instead of degrading gracefully. ### REMEDIAL ACTIONS PLAN & NEXT STEPS We recognize that outages like this have a direct impact on customers’ ability to receive important notifications, complete account tasks, and operate their sites. We are prioritizing the following actions to improve our existing testing, monitoring and certificate management processes: * **Hardening monitoring and certificate infrastructure** * We are refining DNS and certificate configuration across our monitoring domains and strengthening proactive checks to detect and address failed renewals and certificate issues well before expiry. * We are also improving alerting on our monitoring and metrics pipeline. * **Decoupling monitoring from critical customer flows** We are updating services such as outbound email, identity, mobile push, provisioning, and admin policy changes so they no longer depend on metrics publishing to operate. If monitoring becomes unavailable, these services will continue to run and degrade gracefully by dropping or buffering metrics instead of failing customer operations. We apologize to customers impacted during this incident. We are implementing the improvements above to help ensure that similar issues are avoided. Thanks, Atlassian Customer Support
We have successfully mitigated the incident and all affected services are now fully operational. Our teams have verified that normal functionality has been restored across all areas. Thank you for your patience and understanding while we worked to resolve this issue.
We have taken steps to mitigate the issue and are seeing recovery in the affected services. Our teams will continue to closely monitor the situation and are actively working to confirm that all services are fully restored. We will provide further updates as we make additional progress.
500 Errors being experienced by Trello users
4 updates
### Summary On December 12, 2025, between 05:31 UTC and 07:11 UTC, Trello customers experienced errors or slow loading times due to a dependency library upgrade that caused high resource usage. The incident was detected within 1 minute by the automated monitoring system and mitigated by rolling back the version of the dependency library which put Atlassian systems into a known good state. The total time to resolution was 1 hour and 41 minutes. ### **IMPACT** The issue caused service disruption for all Trello customers on December 12, 2025 between 05:31 UTC and 07:11 UTC. Users experienced degraded service between 05:31 UTC and 06:32 UTC, and most users were unable to load Trello between 06:32 UTC to 07:07 UTC. ### **ROOT CAUSE** The issue was caused by an upgrade to a dependency library. As a result, there was an increase in system memory consumption causing startups to take longer than expected. When the startup time exceeded the allowed threshold, the worker process was flagged as unresponsive and restarted as an automated recovery action. Because every worker on the instance failed to startup within the expected time, the instance entered a startup loop that required intervention. ### **REMEDIAL ACTIONS PLAN & NEXT STEPS** We know that outages impact your productivity. While we have several testing and preventive processes in place, we didn’t anticipate that the upgrade would consume more memory, which ultimately caused the startup failure loop and service disruption. We are prioritizing the following improvement actions designed to avoid this issue in the future: * Review our automated detection of unresponsive workers and adjust the detection thresholds during initialization. * Improve monitoring during workers startup and restarts due to unresponsiveness. * Enhance the stricter standard process for confirming memory/other data processing resource requirements for planned updates. We apologize to customers whose services were impacted during this incident; we are taking immediate steps to improve the platform’s performance and availability. Thanks, Atlassian Customer Support
On December 12, 2025, from 05:51 UTC to 08:09 UTC, Trello users may have experienced service disruptions on the web page and mobile apps. The issue has now been resolved, and the service is operating normally for all affected customers.
Our engineering team has implemented a fix, and services have now recovered. We will continue to closely monitor the situation.
We are actively investigating reports of a service disruption affecting Trello users. We'll share updates here as they become available or within the hour.
November 2025(2 incidents)
Multiple Atlassian services experiencing degraded performance
5 updates
### Summary On November 21, 2025, between 13:44 and 15:16 UTC, Trello customers were intermittently unable to view and update data on their boards. Customers also may have experienced issues authenticating with Atlassian products, and creating new GitHub and Slack integrations. The event was triggered by a bug encountered in the software running our edge proxy fleet, which proxies customer traffic to Atlassian cloud services. The changes included the migration of our edge proxy fleet to hosts running an ARM CPU architecture, rather than the AMD64 CPU architecture they had previously been running, which impacted US East customers. The incident was detected within 1 minute by our automated monitoring systems, and mitigated by a scale up of of fleet size, which put Atlassian systems into a known good state. This was followed by a global migration of edge proxy fleet hosts back to AMD64 CPU architecture the following day. ### **IMPACT** During the impact window, US East customers intermittently could not view or update data in Trello. The same underlying issue also impacted our Identity services and integrations with GitHub and Slack, meaning some customers had trouble signing in to Atlassian products or creating new integrations. At the incident’s peak, the incident impacted up to: * 52% of new Trello network connections. * 9% of new GitHub and Slack integrations. * 8% of new Identity network connections. ### **ROOT CAUSE** The issue was caused by a change to CPU architecture from AMD64 to ARM on our edge proxy fleet. This led to a bug that caused these instances to stall under high load, and refuse up to 52% of new connections. As a result, some customers of the products above could not make new connections to Atlassian services, and customers received CloudFront 504 gateway timeout error responses. ### **REMEDIAL ACTIONS PLAN & NEXT STEPS** We know that outages impact your productivity. While we deploy our changes progressively by cloud region to avoid broad impact, on this occasion, our pre-change load testing had not accurately reflected production loads. As part of our response to this incident, and to help prevent recurrence, we rolled back all edge proxy fleets from ARM to AMD64 CPU architecture globally. To minimise the impact of breaking changes to our environments, we plan to implement additional preventative measures such as: * Adding improved load tests into edge proxy fleet deployment pipelines to catch load-related bugs before deployment to production. * Adding alerts to our edge proxy fleet to catch rises in TCP connect times before customer impact. We apologize to customers whose services were impacted during this incident; we are taking steps to help improve the platform’s performance and availability. Thanks, Atlassian Customer Support
The Trello performance degradation has been resolved. Also, some intermittent errors on some of our other products, has also now been resolved. The issue has now been resolved, and the service is operating normally for all affected customers.
The Trello performance degradation affecting some customers has been resolved. A subset of customers may have experienced intermittent errors on some of our other products, but these should also now be resolved too. We'll continue to monitor closely to confirm stability.
Multiple Atlassian services are experiencing degraded performance. We are investigating and will provide an update within the hour.
Multiple Atlassian services are seeing outages and we are investigating the same. We shall keep you posted on the progress in 60 minutes, if not sooner
Trello automations delayed
4 updates
On November 14, 2025 Trello users may have experienced delays on the execution of Trello automations. The issue has now been resolved, and the service is operating normally for all affected customers. Note: All automations identified as impacted by this incident will be automatically rerun, no customer actions are needed to rerun these automations.
We have applied mitigations and continue to see improvements in Trello Automation processing times. However, some users may still experience delays in automation execution. We will provide another update by 20:45 UTC, or sooner if more information becomes available.
We have identified the cause of the issue, and our teams are diligently working on applying mitigations. We are seeing good signs of recovery, however, some users may experience some delays in Trello Automations executing. We’ll share additional updates at 19:45 UTC, or sooner, as more information is available.
We are actively investigating reports of performance degradation affecting Trello automations. We'll share updates here as more information is available.
October 2025(1 incident)
Atlassian Cloud Services impacted
25 updates
### Postmortem publish date: Nov 19th, 2025 ### Summary All dates and times below are in UTC unless stated otherwise. Customers utilizing Atlassian products experienced elevated error rates and degraded performance between Oct 20, 2025 06:48 and Oct 21, 2025 04:05. The service disruptions were triggered due to an [AWS DynamoDB outage](https://aws.amazon.com/message/101925/#:~:text=1%3A50%20PM.-,DynamoDB,-Between%2011%3A48) and further affected by subsequent failures in [AWS EC2](https://aws.amazon.com/message/101925/#:~:text=service%20disruption%20event.-,Amazon%20EC2,-Between%2011%3A48) and [AWS Network Load Balancer](https://aws.amazon.com/message/101925/#:~:text=service%20disruption%20event.-,Amazon%20EC2,-Between%2011%3A48) within the us-east-1 region. The incident started at Oct 20, 2025 06:48 and was detected within six minutes by our automated monitoring systems. Our teams worked to restore all core services by Oct 21, 2025 04:05. Final cleanup of backlogged processes and minor issues was completed on Oct 22, 2025. We recognize the critical role our products play in your daily operations, and we offer our sincere apologies for any impact this incident had on your teams. We are taking immediate steps to enhance the reliability and performance of our services, so that you continue to receive the standard of service you have come to trust. ### IMPACT Before examining product-level impacts, it's helpful to understand Atlassian's service topology and internal dependencies. Products such as Jira and Confluence are deployed across multiple AWS regions. The data for each tenant is stored and processed exclusively within its designated host region. This design is intentional and represents the desired operational state, as it limits the impact of any regional outage strictly to tenants in-region, in this case us-east-1. While in-scope application data is pinned to the region selected by the customer, there are times when systems need to call other internal services that may be based in a different region. If a problem occurs in the main region where these services operate, systems are designed to automatically fail over to a backup region, usually within three minutes. However, if unexpected issues arise during this failover, it can take longer to restore services. In rare cases, this could affect customers in more than one region. It’s important to note that all in-scope application data for supported products is pinned according to a customer’s chosen region. **Jira** Between Oct 20, 2025 06:48 and Oct 20, 2025 20:00, customers with tenants hosted in the us-east-1 region experienced increased error rates when accessing core entities such as Issues, Boards, and Backlogs. This disruption was caused by AWS's inability to allocate AWS EC2 instances and elevated errors in AWS Network Load Balancer \(NLB\). During this window, users may also have observed intermittent timeouts, slow page loads, and failures when performing operations like creating or updating issues, loading board views, and executing workflow transitions. Between Oct 20, 2025 08:36 and Oct 20, 2025 09:23, customers across all regions experienced elevated failure rates when attempting to load Jira pages. This disruption was caused by the regional frontend service entering an unhealthy state during this specific time interval. Normally, the frontend service connects to the primary AWS DynamoDB instance located in the us-east-1 to retrieve the most recent configuration data necessary for proper operation. Additionally, the service is designed with a fallback mechanism that references static configuration data in the event that the primary database becomes inaccessible. Unfortunately, a latent bug existed in the local fallback path. When the frontend service nodes restarted, they were unable to load critical operational configuration data from primary or fallback sources, leading to the observed failures experienced by customers. Between Oct 20, 2025 06:48 and Oct 21, 2025 06:30, customers experienced significant delays and missing Jira in-app notifications across all regions. The notification ingestion service, which is hosted exclusively in us-east-1, exhibited an increased failure rate when processing notification messages due to AWS EC2 and NLB issues. This issue resulted in notifications being delayed - and in some cases, not delivered at all - to users worldwide. **Jira Service Management \(JSM\)** JSM was impacted similarly to Jira above, with the same timeframes and for the same reasons. Between Oct 20, 2025 08:36 and Oct 20, 2025 09:23, customers across all regions experienced significantly elevated failure rates when attempting to load JSM pages. This affected all JSM experiences including the Help Centre, Portal, Queues, Work Items, Operations, and Alerts. **Confluence** Between Oct 20, 2025 06:48 and Oct 21, 2025 02:45, customers using Confluence in the us-east-1 region experienced elevated failure rates when performing common operations such as editing pages or adding comments. The primary cause of this service degradation was the system's inability to auto-scale due to AWS EC2 issues to manage peak traffic load effectively. Though the AWS outage ended at Oct 20, 21:09, a subset of customers continued to experience failures as some Confluence web server nodes across multiple clusters remained in an unhealthy state. This was ultimately mitigated by recycling the affected nodes. To protect our systems while AWS recovered, we made a deliberate decision to enable node termination protection. This action successfully preserved our server capacity but, as a trade-off, it extended the time required for a full recovery once AWS services were restored. **Automation** Between Oct 20, 2025 06:55 and Oct 20, 2025 23:59, automation customers whose rules are processed in us-east-1 experienced delays of up to 23 hours in rule execution. During this window, some events triggering rule executions were processed out of order because they arrived later during backlog processing. This caused potential inconsistencies in workflow executions, as rules were run in the order events were received, not when the action causing the event occurred. Additionally, some rule actions failed because they depend on first-party and third-party systems, which were also affected by the AWS outage. Customers can see most of these failures in their audit logs; however, a few updates were not logged due to the nature of the outage. By Oct 21, 2025 5:30, the backlog of rule runs in us-east-1 was cleared. Although most of these delayed rules were successfully handled, there were some additional replays of events to ensure completeness. Our investigation confirmed that a few events may never have triggered their associated rules due to the outage. Between Oct 20, 2025 06:55 and Oct 20, 2025 11:20, all non-us-east-1 regional automation services experienced delays of up to 4 hours in rule execution. This was caused by an upstream service that was unable to deliver events as expected. The delivery service encountered a failure due to a cross-region dependency call to a service hosted in the us-east-1 region. Because of this dependency issue, the delivery service was unable to successfully deliver events throughout this time frame, resulting in customer-defined rules not being executed in a timely manner. **Bitbucket and Pipelines** Between Oct 20, 2025 06:48 and Oct 20, 2025 09:33, Bitbucket experienced intermittent unavailability across core services. During this period, users faced increased error rates and latency when signing in, navigating repositories, and performing essential actions such as creating, updating, or approving pull requests. The primary cause was an AWS DynamoDB outage that impacted downstream services. Between Oct 20, 2025 06:48 and Oct 20, 2025 22:46, numerous Bitbucket Pipeline steps failed to start, stalled mid-execution, or experienced significant queueing delays. Impact varied, with partial recoveries followed by degradation as downstream components re-synchronized. The primary cause was an AWS DynamoDB outage, compounded by instability in AWS EC2 instance availability and AWS Network Load Balancers. Furthermore, Bitbucket Pipelines continued to experience a low but persistent rate of step timeouts and scheduling errors due to AWS bare-metal capacity shortages in select availability zones. Atlassian coordinated with AWS to provision additional bare-metal hosts and addressed a significant backlog of pending pods, successfully restoring services by 01:30 on Oct 21, 2025. **Trello** Between Oct 20, 2025 06:48 and Oct 20, 2025 15:25, users of Trello experienced widespread service degradation and intermittent failures due to upstream AWS issues affecting multiple components, including AWS DynamoDB and subsequent AWS EC2 capacity constraints. During this period, customers reported elevated error rates when loading boards, opening cards, adding comments or attachments. **Login** Between Oct 20, 2025 06:48 and Oct 20, 2025 09:30, a small subset of users experienced failures when attempting to initiate new login sessions using SAML tokens. This resulted in an inability for those users to access Atlassian products during that time period. However, users who already had valid active sessions were not affected by this issue and continued to have uninterrupted access. The issue impacted all regions globally because regional identity services relied on a write replica located in the us-east-1 region to synchronize profile data. When the primary region became unavailable, the failover to a secondary database in another region failed, which delayed recovery. This failover defect has since been addressed. **Statuspage** Between Oct 20, 2025 06:48 and Oct 20, 2025 09:30, Statuspage customers who were not already logged in to the management portal were unable to log in to create or update incident statuses. This impact was restricted only to users who were not already logged in at the time. The root cause was the same as described in the Login section above, and it was resolved by the same remediation steps. ### REMEDIAL ACTION PLAN & NEXT STEPS We have completed the following critical actions designed to help prevent cross-region impact from similar issues: * Resolved the code defect in the fallback option to ensure that Jira Frontend Services in other regions remain unaffected during a region-wide outage. * Fixed the issue that prevented timely failover of the identity service which impacted new login sessions. * Resolved the code defect so that delivery services in unaffected regions remain operational during region-wide outages. Additionally, we are prioritizing the following improvement actions: * Implement mitigation strategies to strengthen resilience against region-wide outages in the notification ingestion service. Although disruptions to our cloud services are sometimes unavoidable during outages of the underlying cloud provider, we continuously evaluate and improve test coverage to strengthen resilience of our cloud services against these issues. We recognize the critical importance of our products to your daily operations and overall productivity, and we extend our sincere apologies for any disruptions this incident may have caused your teams. If you were impacted and require additional details for internal post-incident reviews, please reach out to your Atlassian support representative with affected timeframes and tenant identifiers so we can correlate logs and provide guidance. Thanks, Atlassian Customer Support
Our team is now able to see full recovery across the vast majority of Atlassian products. We are aware of some ongoing issues with specific components such as migrations and JSM virtual service agents, and our team is continuing to investigate with urgency. We apologise for the inconvenience that this incident has caused and we will provide further information when the Post Incident Investigation has been completed.
The issue relating to the Atlassian Support portal displaying a message to customers to use our temporary support channel has now been resolved. The Atlassian Support portal is fully functional for any ongoing support issues. With regards to other Atlassian products, we continue to see recovery continuing across all impacted products and our teams are continuing to monitor as the recovery continues. We will provide further update on our recovery status within two hours.
We continue to see recovery progressing across all impacted products as backlogged items continue to be processed. The Atlassian Support portal is currently displaying a message directing customers to our temporary support channel. Please note that our support portal is currently fully functional for those attempting to raise requests. We are continuing to look into this alert to remove this message. We will provide further update on our recovery status in two hours.
Our team is now seeing recovery across all impacted Atlassian products. We are continuing to monitor for individual products that may still be processing backlogged items now that services are restored. The Atlassian Support portal is currently still displaying a message directing customers to our temporary support channel. Please note that our support portal is currently fully functional for those attempting to raise requests. We are continuing to look into this alert to remove this message. We will provide further update on our recovery status in one hour.
Our teams are continuing to monitor the recovery of systems across Atlassian products. This update is to inform that the Atlassian Support portal is fully operational at this time for customers that wish to contact support.
Monitoring - We've started seeing continued product experience improvement. While we still have a backlog of event processing, we are seeing improvements in systems operational capabilities across all products. We estimate a significant improvement with the next few hours and will continue to monitor the health of AWS services and the effects on Atlassian customers. We appreciate your continued patience and remain committed to full resolution as we work through this situation. We will post our next update in two hours.
There have been no changes since our last update. We will provide our next updated by 9:00PM UTC or sooner as new information becomes available. We are currently aware of an ongoing incident impacting Atlassian Cloud services due to an outage with our public cloud provider, AWS. We understand the impact this issue is having on your operations and want to assure you that resolving this matter is our highest priority and we are closely monitoring the health of AWS services. While we do not have a definitive ETA at this time, we remain committed to full resolution and deeply appreciate your patience as we work through this situation.
We are currently aware of an ongoing incident impacting Atlassian Cloud services due to an outage with our public cloud provider, AWS. We understand the impact this issue is having on your operations and want to assure you that resolving this matter is our highest priority and we are closely monitoring the health of AWS services. While we do not have a definitive ETA at this time, we remain committed to full resolution and deeply appreciate your patience as we work through this situation. We will continue to provide updates every hour or sooner as new information becomes available.
Update - We understand the impact this issue is having on your operations and want to assure you that resolving this matter is our highest priority. Our public cloud provider is actively working to mitigate this issue with urgency. While we do not have a definitive ETA at this time, we remain committed to full resolution and deeply appreciate your patience as we work through this situation. We will continue to provide updates every hour or sooner as new information becomes available.
Update - Thank you for your continued patience. We understand the impact this issue is having on your operations and want to assure you that resolving this matter is our highest priority. Our public cloud provider is still actively working to mitigate this issue with urgency and while we do not have a definitive ETA at this time, we remain committed to full resolution and deeply appreciate your patience as we work through this situation. We will be providing hourly updates on this issue.
Update - Thank you for your continued patience. We understand the impact this issue is having on your operations and want to assure you that resolving this matter is our highest priority. Our public cloud provider is still actively working to mitigate this issue with urgency and while we do not have a definitive ETA at this time, we remain committed to full resolution and deeply appreciate your patience as we work through this situation. We will be providing hourly updates on this issue.
Update - We understand the impact this issue is having on your operations and want to assure you that resolving this matter is our highest priority. Our public cloud provider is actively working to mitigate this issue with urgency. While we do not have a definitive ETA at this time, we remain committed to full resolution and deeply appreciate your patience as we work through this situation. We will continue to provide updates every hour or sooner as new information becomes available.
Update - We understand the impact this issue is having on your operations and want to assure you that resolving this matter is our highest priority. Our public cloud provider is actively working to mitigate this issue with urgency. While we do not have a definitive ETA at this time, we remain committed to full resolution and deeply appreciate your patience as we work through this situation. We will continue to provide updates every hour or sooner as new information becomes available.
We understand your pain and mitigating or fixing this issue is of utmost importance. Our public cloud provider is actively working to mitigate this issue on priority. We have been seeing partial operational success. We appreciate your patience and will continue to provide updates every hour or sooner.
Our public cloud provider is working to mitigate this issue quickly. We are seeing some early positive indicators and are continuing to monitor. We appreciate your patience and will continue to provide updates every hour or sooner.
Our public cloud provider is working to mitigate this issue quickly. We are seeing some early positive indicators and are continuing to monitor. We appreciate your patience and will continue to provide updates every hour or sooner.
Our public cloud provider is working to mitigate this issue quickly. We are seeing some early positive indicators and are continuing to monitor. We appreciate your patience and will continue to provide updates every hour or sooner.
Atlassian team is actively engaged and continues to work with our public cloud provider to mitigate this issue at the earliest. We are starting to see partial operations succeed. We appreciate your patience. We shall continue to share updates every hour, if not sooner.
We continue to work with our public cloud provider towards mitigating the issue at the earliest. We appreciate your patience. We shall continue to share updates every hour, if not sooner.
We understand that our public cloud provider has identified the cause of the issue. We are starting to see some recovery and is working towards mitigation. We appreciate your patience. We shall continue to share updates every hour, if not sooner.
We are experiencing an outage due to some issue at the end of our public cloud provider. We are working closely with them to get this resolved or mitigated as quickly as possible. ETA of the same is not know at the moment. We shall continue to share updates every hour, if not sooner.
We are experiencing an outage due to some issue at the end of our public cloud provider. We are working closely with them to get this resolved or mitigated as quickly as possible. ETA of the same is not know at the moment. We shall continue to share updates every hour, if not sooner.
Atlassian Cloud services are impacted and we are aware that our customers might not be able to create support tickets. Our teams are actively investigating the same. We shall keep you informed of the progress every hour.
We have noticed that Atlassian Cloud services are impacted and our teams are actively investigating the same. We shall keep you informed of the progress every hour.
September 2025(1 incident)
Degraded performance of Trello
2 updates
On Thursday, September 11, 2025 1:00PM UTC, Trello users may have experienced performance degradation. The issues has now been resolved, and the service is operating normally for all affected customers.
The performance degradation of Trello has been resolved, and services are now operating normally for all affected customers. We'll continue to monitor performance closely to confirm stability.
August 2025(1 incident)
Trello Connectivity Problems
3 updates
This incident has been resolved.
We are now seeing recovery and monitoring.
We're receiving reports of connectivity problems while accessing Trello. You may see a "You are disconnected from Trello message" in the bottom left corner of your window notifying you that you've been disconnected. We're currently investigating the cause, and will post updates here as we determine it.
July 2025(2 incidents)
Degraded experience adding and accessing media attachments
2 updates
We have resolved this incident.
We are currently investigating this issue.
Users experiencing issues with Email to Board failing
3 updates
Our team has been able to identify a faulty configuration internally that was causing the errors relating to Email to Board functionality in Trello. The faulty configuration has now been reverted and Email to Board functionality should now be working as expected. For those that experienced Email to Board failures these emails can now be re-submitted if required to create the cards in Trello. We apologise again for any inconvenience caused by this issue.
Our team is continuing to investigate the cause of errors relating to Email to Board actions in Trello. For those continuing to experience issues with Email to Board we encourage users to manually create their cards in Trello directly in the meantime. We apologise for any inconvenience caused and we will provide further updates as the investigation continues.
We are aware of some users experiencing issues where Email to Board functionality is failing with an error message similar to 'This message could not be delivered due to a recipient error. Please try again later.' Our team is investigating with urgency and will provide an update as soon as possible.
June 2025(3 incidents)
Issues affecting user syncing, Atlassian Administration
4 updates
Between 07:40 UTC and 10:31 UTC, we experienced issues affecting user syncing in Atassian Administration. This affected Confluence, Jira Work Management, Jira Service Management, Jira, Trello, and Guard. The issue has been resolved and the service is operating normally.
We have identified the root cause of the issue and have mitigated the problem. We are now monitoring closely.
We continue to work on resolving the user syncing functionality for Confluence, Jira Work Management, Jira Service Management, Jira, Trello, and Guard. We have identified the root cause and expect recovery shortly.
We are investigating reports of errors loading Users and Groups pages in Atlassian Administration, and errors affecting user IDP syncing. We will provide more details once we identify the root cause.
Customers may experience delays receiving emails
2 updates
Between 2025-06-04 14:11 UTC to 20:18 UTC, we experienced delays in delivering emails for Confluence, Jira Work Management, Jira Service Management, Jira, Trello, Atlassian Bitbucket, Guard, Jira Align, Jira Product Discovery, Atlas, Compass. The issue has been resolved and the service is operating normally.
We were experiencing cases of degraded performance for outgoing emails from Confluence, Jira Work Management, Jira Service Management, Jira, Trello, Atlassian Bitbucket, Guard, Jira Align, Jira Product Discovery, Atlas and Compass Cloud customers. The system is recovering and mail is being processed normally as of 16:45 UTC. We will continue to monitor system performance and will provide more details within the next hour.
The search bar functionality in Trello is not working properly
4 updates
On June 2nd, Trello's search bar functionality was not working correctly. The issue has now been resolved, and the service operates normally for all affected customers.
The issues causing Trello’s search bar functionality to malfunction have been resolved, and services are now operating normally for all affected customers. We will monitor it closely to ensure stability.
We are currently experiencing an issue with Trello, as the search bar functionality is not working properly. We have identified the issue and are in the process of mitigating the impact. We'll keep you posted with further updates.
We are currently experiencing an issue with Trello, as the search bar functionality is not working properly. Our team is diligently working to restore services as quickly as possible. We will keep you updated with further developments.
May 2025(2 incidents)
Trello was temporarily inaccessible
3 updates
### **SUMMARY** On May 15, 2025, between 13:55 and 14:18 UTC, Atlassian customers using the Trello product experienced errors or slow loading times when attempting to view their cards and boards. The event was triggered by a database plan cache expiring and high resource usage caused by subsequent database query planning operations. The particular database shard that was impacted held data that was required for every card load. The incident was detected within two minutes by the automated monitoring system and mitigated by increasing resources available to the affected database shard, which put Atlassian systems into a known good state. The total time to resolution was about 23 minutes. ### **IMPACT** The overall impact was between May 15, 2025, 13:55 and May 15, 2025, 14:18 UTC on the Trello product. The incident caused service disruption for all Trello customers. ### **ROOT CAUSE** The issue was caused by a query plan expiring from the database cache, which caused incoming queries to go through a replanning operation. These queries had multiple plans that could satisfy them, and depending on the size of the query, one plan might be significantly more efficient than another. This caused the query planner to perform a great many more replanning operations than usual, which consumed all of the CPU on the server for a brief moment. Once the CPU was consumed, the planning operations themselves began taking too long and therefore required constant replanning in an effort to find more efficient options. This negative feedback loop could not be broken without intervention. ### **REMEDIAL ACTIONS PLAN & NEXT STEPS** We know that outages impact your productivity. While we have a number of testing and preventative processes in place, this specific issue wasn’t identified because it would only occur under very distinct conditions, including the amount of load and the order of database queries. We are prioritizing the following improvement actions to avoid repeating this type of incident: * Review our capacity planning thresholds and ensure that all shards have sufficient overhead to handle unexpected load. * Improve query planner performance by: * Implement hinting for known problematic query shapes to circumvent the query planner. * Investigate long-term generalized solutions to prevent query planner thrashing. Furthermore, we are prioritizing the following additional measures to reduce the impact of any future incidents: * Analyze and reduce single points of failure for loading Trello boards and cards. We apologize to customers whose services were impacted during this incident; we are taking immediate steps to improve the platform’s performance and availability. Thanks, Atlassian Customer Support
A fix has been implemented, and the issue is now resolved.
We are aware that Trello was temporarily inaccessible for all customers. The system has already been recovered, and we are monitoring the situation.
Trello is slow or unavailable for some users
4 updates
### **SUMMARY** On May 5, 2025, between 2:08 p.m. and 4:29 p.m. UTC, some Atlassian customers using Trello were unable to view their boards or cards. The event was triggered by an unexpected error encountered by our infrastructure management tools, which resulted in an incorrect DNS configuration being deployed to a portion of our database. The incident was detected within four minutes by automated monitoring systems and mitigated by identifying the faulty portion of the database and performing a failover, which put Atlassian systems into a known good state. The total time to resolution was about two hours and 21 minutes. ### **IMPACT** The overall impact was on the Trello product on May 5, 2025, between 2:08 p.m. and 4:29 p.m. UTC. The incident caused service disruption to Trello customers whose accounts and boards contained or referenced data on the affected shard of our database. Additionally, some Trello customers would have experienced a service disruption due to our use of load-shedding tools during the incident to strategically block portions of our traffic to aid in recovery. ### **ROOT CAUSE** The day before the incident, on May 4, our infrastructure management tooling encountered an unexpected error when attempting to fetch the networking metadata on a particular host. This led to the host, which was a member of our database cluster, to incorrectly apply the default Operating System DNS configuration. This DNS configuration was not able to resolve internal domains, which led to a partial failure state of the node. The database continued to function normally and there was no immediate customer impact but in the background this incorrect DNS configuration led to the slow buildup of database sessions. These database sessions are usually short-lived and automatically expire when no longer needed, but the DNS misconfiguration prevented this automatic expiration. The database sessions eventually grew to the default maximum on this particular shard. At that point, the shard was unable to generate new sessions, which are required for all basic operations, and the Trello product began experiencing elevated error rates. ### **REMEDIAL ACTIONS PLAN & NEXT STEPS** We know that outages impact your productivity. While we have a number of testing and preventative processes in place, this specific issue wasn’t identified due to the isolated nature of the database session resource and monitoring gaps around this resource and around DNS resolution. We are prioritizing the following improvement actions designed to avoid repeating this type of incident: * Update our infrastructure management tool to use a safe fall-back DNS configuration in the case of unexpected errors. * Expand existing DNS monitoring to include the resolution of internal domains. * Expand existing database session count monitoring to include all database node types. Furthermore, we are prioritizing the following additional measures to reduce the duration of any future incidents: * Evaluate our incident response process to identify actions that can be streamlined for quicker resolution. We apologize to customers whose services were impacted during this incident; we are taking steps designed to improve the platform’s performance and availability. Thanks, Atlassian Customer Support
On May 5th, 2025 we identified a degradation for Trello. Trello is now back online and no further impact has been observed.
We have identified and mitigated the issue causing Trello to be slow or unavailable for some users. We expect API traffic to return to normal within the next 30 minutes. We are now monitoring closely. We will update within the next 30 minutes.
We've noticed that Trello is slow or unavailable for some users. This will be present in both the web and mobile apps. Our engineering team is actively investigating this incident and working to bring Trello back up to speed as quickly as possible. We'll keep you posted with further updates on this page.
April 2025(1 incident)
Trello is slow or unavailable
3 updates
### **SUMMARY** On April 25, 2025, between 18:18 and 18:33 UTC, Atlassian customers using Trello may have experienced service interruptions. The event was triggered by temporarily reduced capacity following a rollback deployment, with insufficient nodes to handle the load. Automated monitoring systems detected the incident within one minute and mitigated it by scaling up the deployment, which put Atlassian systems into a known-good state. The total time to resolution was about 15 minutes. ### **IMPACT** The overall impact was on April 25, 2025, between 18:18 and 18:33 UTC, on Trello. The incident caused service disruption to Trello users, resulting in reduced functionality, slower response times, and errors when performing key actions such as loading boards and cards. ### **ROOT CAUSE** The root cause of the incident was a failure to scale our nodes to optimal capacity caused by a release rollback. If an issue is found during a deployment, we can roll back to a previous release. In this case, a rollback was executed to a previous release that had already undergone a scaling-down process. As that rollback happened, more compute nodes needed to be available to handle the high traffic. ### **REMEDIAL ACTIONS PLAN & NEXT STEPS** We know that outages impact your productivity and strive to avoid incidents like these. We are prioritizing the following efforts as next steps: * Improvements to rollback tooling including UX upgrades and pre-scaling * Conduct an updated incident response training focused on rollback tooling and best practices We apologize to customers whose services were impacted during this incident; we are taking steps designed to improve the platform’s performance and availability. Thanks, Atlassian Customer Support
Between 18:18 UTC to 18:33 UTC, we experienced an outage for Trello. The issue has been resolved and the service is operating normally. Our teams are investigating and will publish the root cause as soon as available.
Trello was slow or unavailable. Our engineering team is actively investigating this incident to determine root cause. Users affected by this incident may have noticeed that Trello was slow or completely unavailable in both the web and mobile apps. Trello operations have recovered. We will update this page as we have additional information.
March 2025(4 incidents)
Degraded performance in Trello
6 updates
Following a period of monitoring, we've confirmed that the issue with degraded performance in Trello has been resolved. For any further issues please contact our Support team.
Trello is back up and running and we're not seeing any more performance issues but we'll continue to monitor this. If you're still experiencing any performance issues, please reach out to us through our support channels for further assistance
We're still looking into issues with degraded performance in Trello and our team are working to get Trello back up and running as quickly as possible. We'll post more updates on this shortly.
We're continuing to investigate issues causing degraded performance in Trello for some Trello customers. We'll post more updates on this shortly.
Our team have implemented a fix and we're seeing performance improve, we'll continue to monitor this to make sure everything is working as normal.
We're currently investigating an issue that's causing degraded performance in Trello for some customers and our team are working to get Trello back up and running as quickly as possible.
Trello Realtime Updates are degraded
2 updates
We've identified this issue and released a fix so realtime updates are back up and running in Trello. We appreciate your patience and understanding while we worked on this.
We're currently investigating an issue causing degraded performance for Trello's realtime updates for some customers. Our team are actively working to bring Trello back up to speed as quickly as possible.
Trello is slow
4 updates
The changes we applied to our infrastructure have resulted in very good results on our side, and we’re no longer noticing issues that would result in constant disconnections or degraded services for Trello. We appreciate your understanding while we worked on this.
We’ve already identified the issue and are working on a possible fix in our infrastructure for all connection problems some users have been experiencing. We’ll share a new update soon!
We're continuing to investigate this issue with connections to Trello being dropped. Our engineering are working to bring Trello back up to speed as quickly as possible.
We are investigating an issue related to connections being dropped. Some Trello users are experiencing slower updates and disconnection messages. Our engineering team is actively working to bring Trello back up to speed as quickly as possible.
Trello Realtime Updates are degraded
5 updates
Between to , we experienced realtime update functionality degraded performance for Trello. We have mitigated the problem and are investigating the root cause. The issue has been resolved and the service is operating normally.
Between to , we experienced realtime update functionality degraded performance for Trello. We have mitigated the problem and are investigating the root cause. We are now monitoring closely.
We are continuing our investigation on realtime updates that is impacting all Trello Cloud customers. We will provide more details within the next hour.
We are still investigating realtime update errors that is impacting most Trello Cloud customers. We will provide more details within the next hour.
We are investigating cases of degraded performance for SOME Trello Cloud customers. We will provide more details within the next hour.
February 2025(1 incident)
The Trello Desktop app is re-directing users to the home page
5 updates
Following a period of monitoring, we have confirmed that the issue affecting the MacOS and Windows Desktop Apps is resolved. Please update to the latest version for the latest fix. For any further issues please contact our Support team. Thank you for your patience.
Our engineering team has successfully released a fix for both the MacOS and Windows Desktop Apps. To benefit from this fix, users are required to update to the latest version of the app. If you continue to experience any issues, particularly with redirection, please reach out to us through our Support channels for further assistance. Thank you for your patience and understanding as we worked to resolve this issue.
We're still working on a fix for the Trello Desktop. We'll provide a new update soon.
We're still working on a fix for the Trello Desktop. We'll provide a new update soon.
We've identified an issue that's causing users to be re-directed to the Home page when using the Trello Desktop app. We'll provide a new update soon.
December 2024(1 incident)
Users aren't able to create new repeats or run existing ones when using the Card Repeater Power-Up
5 updates
A fix was released for the Card Repeater Power-up, and the repeats and new ones are working as intended.
Our recent change hasn’t fixed the existing issue with the Card Repeater Power-Up. We’ll keep investigating the problem, and a new update will be provided soon.
A fix has been implemented and we're currently monitoring it.
We've identified the issue that's causing issues with the repeats, and we're working now on a fix.
We’re currently investigating reports that users can’t create new repeats while using the Card Repeater and that existing repeats aren’t being executed.
October 2024(1 incident)
Boards are demonstrating slow performance for Workspace Guests
8 updates
On October 15th, we identified that some boards were performing slowly for Workspace Guests. A fix has been released to all affected users, and the slow performance issue is fixed. No further impact has been observed.
On October 15th, we identified that some boards were performing slowly for Workspace Guests. We have taken action to mitigate this issue and we will continue to monitor for the next 30 minutes before marking the incident as resolved.
We've identified the issue that caused the issue for slow performance for boards when loading them and opening cards for Workspace Guests, and we're now working on a fix for the problem and will provide a new update here soon.
We’re continuing to investigating issues with boards that are experiencing slow performance when loading them and opening cards for Workspace Guests, and we will provide updates here soon
We are continuing to investigate issues with some large boards that are experiencing poor performance when loading them and opening cards, and will provide updates as we learn more.
We are continuing to investigate issues with some large boards that are experiencing poor performance when loading them and opening cards, and will provide updates as we learn more.
We are continuing to investigate issues with some large boards that are experiencing poor performance when loading them and opening cards, and will provide updates as we learn more.
We are investigating issues with some large boards that are experiencing poor performance when loading them and opening cards, and we will provide updates here soon.
September 2024(2 incidents)
Email to Board feature isn't working as expected for some customers
5 updates
On September 25th, we identified a temporary outage with the Email-to-Board feature. The affected feature is now back online, and no further impact has been observed.
On Sept 25 we identified that our email-to-board feature was failing for some users. We have taken action to mitigate this issue and we will continue to monitor for the next 30 minutes before marking the incident as resolved.
We've identified the issue that caused the issues with the Email-to-Board feature not working correctly, and we're now working on a fix for the problem and will provide a new update here soon.
We are continuing to investigate issues with Email-to-Board feature not working for some customers and will provide updates as we learn more.
We are investigating issues with Email-to-Board feature not working for some customers and will provide updates here soon.
Users are experiencing reCaptcha errors while signing up
3 updates
This issue has been resolved.
We have identified the root cause and the issue appears to be resolved.
Users attempting to sign up are encountering reCaptcha errors that are preventing a successful signup.
July 2024(3 incidents)
Issue with search and moving cards for some users
4 updates
We previously identified the issue with moving cards and searching for cards in Trello. The issue has been resolved and Trello is operating normally.
We’ve identified the issue causing problems with moving and searching for cards on Trello. We’re currently working on implementing a fix, and a new update will be shared soon.
We are investigating reports of intermittent errors for some Trello customers when moving cards between boards or searching for cards. We will provide more details once we identify the root cause
We are investigating reports of intermittent errors for some Trello customers when moving cards between boards or searching for cards. We will provide more details once we identify the root cause.
Some users may experience delays in receiving email notifications
3 updates
Between 12:00am 9th July to 08:00am 10th July, we experienced email deliverability issues for some recipient domains for Confluence, Jira Work Management, Jira Service Management, Jira, Trello, Atlassian Bitbucket, and Jira Product Discovery. The issue has been resolved and future emails will flow normally.
We continue to work on resolving the Email Notifications for Confluence, Jira Work Management, Jira Service Management, Jira, Trello, Atlassian Bitbucket, and Jira Product Discovery. We have identified the root cause.
We are investigating reports of intermittent errors whilst sending Email Notifications for Confluence, Jira Work Management, Jira Service Management, Jira, Trello, and Jira Product Discovery Cloud customers. We will provide more details once we identify the root cause.
Some products are hard down
3 updates
Between 03-07-2024 20:08 UTC to 03-07-2024 20:31 UTC, we experienced downtime for Trello. The issue has been resolved and the service is operating normally.
We have mitigated the problem and continue looking into the root cause. The outage was between 8:08pm 03/07 UTC - 08:31pm 03/07 UTC We are now monitoring closely.
We are investigating an issue with that is impacting Atlassian, Atlassian Partners, Atlassian Support, Confluence, Jira Work Management, Jira Service Management, Jira, Opsgenie, Atlassian Developer, Atlassian (deprecated), Trello, Atlassian Bitbucket, Guard, Jira Align, Jira Product Discovery, Atlas, Atlassian Analytics, and Rovo Cloud customers. We will provide more details within the next hour.
June 2024(3 incidents)
Issues with the Jira powerup
2 updates
The issues with the Jira powerup have been resolved!
Users of the Jira powerup for Trello may be experiencing errors when installing and using the powerup on their boards. The team is working to resolve this issue as quickly as possible.
Intermittent error accessing content
3 updates
Between 2024-06-20 22:04 UTC to 2024-06-20 22:28 UTC, we experienced intermittent issue for users to access the services for some Atlassian Cloud customers. The issue has been resolved and the service is operating normally.
We have identified the root cause of the intermittent errors and have mitigated the problem. We are now monitoring closely.
We are investigating an intermittent issue with accessing Atlassian Cloud services that is impacting some Atlassian Cloud customers. We will provide more details once we identify the root cause.
Error responses across multiple Cloud products
3 updates
### Summary On June 3rd, between 09:43pm and 10:58 pm UTC, Atlassian customers using multiple product\(s\) were unable to access their services. The event was triggered by a change to the infrastructure API Gateway, which is responsible for routing the traffic to the correct application backends. The incident was detected by the automated monitoring system within five minutes and mitigated by correcting a faulty release feature flag, which put Atlassian systems into a known good state. The first communications were published on the Statuspage at 11:11pm UTC. The total time to resolution was about 75 minutes. ### **IMPACT** The overall impact was between 09:43pm and 10:17pm UTC, with the system initially in a degraded state, followed by a total outage between 10:17pm and 10:58pm UTC. _The Incident caused service disruption to customers in all regions and affected the following products:_ * Jira Software * Jira Service Management * Jira Work Management * Jira Product Discovery * Jira Align * Confluence * Trello * Bitbucket * Opsgenie * Compass ### **ROOT CAUSE** A policy used in the infrastructure API gateway was being updated in production via a feature flag. The combination of an erroneous value entered in a feature flag, and a bug in the code resulted in the API Gateway not processing any traffic. This created a total outage, where all users started receiving 5XX errors for most Atlassian products. Once the problem was identified and the feature flag updated to the correct values, all services started seeing recovery immediately. ### **REMEDIAL ACTIONS PLAN & NEXT STEPS** We know that outages impact your productivity. While we have several testing and preventative processes in place, this specific issue wasn’t identified because the change did not go through our regular release process and instead was incorrectly applied through a feature flag. We are prioritizing the following improvement actions to avoid repeating this type of incident: * Prevent high-risk feature flags from being used in production * Improve the policy changes testing * Enforcing longer soak time for policy changes * Any feature flags should go through progressive rollouts to minimize broad impact * Review the infrastructure feature flags to ensure they all have appropriate defaults * Improve our processes and internal tooling to provide faster communications to our customers We apologize to customers whose services were affected by this incident and are taking immediate steps to address the above gaps. Thanks, Atlassian Customer Support
Between 22:18 UTC to 22:56 UTC, we experienced errors for multiple Cloud products. The issue has been resolved and the service is operating normally.
We are investigating an issue with error responses for some Cloud customers across multiple products. We have identified the root cause and expect recovery shortly.
May 2024(1 incident)
Card Repeater Power-Up failing to repeat
2 updates
Our engineering team has resolved this issue, and the Card Repeater Power-Up is working again. Cards that missed their scheduled repeat will run at their next scheduled time.
Our engineering team identified an incident affecting the Card Repeater Power-Up service which caused an issue while repeating cards and we're currently investigating.
April 2024(2 incidents)
Managing workspace members is failing for some workspace administrators
3 updates
A fix has been deployed. Thank you for your patience.
We have identified the root cause and are working towards releasing a fix.
We are currently investigating this issue.
Admin Portal Feature Access Issue
2 updates
Between 6:30 AM UTC to 9:50 AM UTC, we experienced failures in accessing some features from the Admin Portal. The issue has been resolved and the service is operating normally.
We are investigating an issue causing failures in accessing some features from the Admin Portal, which is impacting some of our Cloud customers. We have identified the root cause and anticipate recovery shortly.
March 2024(1 incident)
Some paginated queries in Forge hosted storage kept repeating the last page
1 update
We have identified and resolved a problem with Forge hosted storage, where some paginated queries kept repeating the last page. The incident was detected by our internal monitoring and was resolved quickly after detection by reverting the deployment. Activating changes recently made to the query cursors for paginated queries introduced a bug that impacted some apps. A small number of requests were impacted over a 16-minute window, while the incident lasted. Timeline: - 25/Mar/24 10:08 p.m. UTC - Impact started, when the changes were deployed to production - 25/Mar/24 10:09 p.m. UTC - Incident was detected - 25/Mar/24 10:24 p.m. UTC - Incident was resolved and impact ended The impact of this incident has been completely mitigated and our monitoring tools confirm that query operations are back to the pre-incident behaviour. We have also resolved the underlying bug and deployed the fix to production, completely eliminating the cause of this incident. We apologise for any inconvenience this may have caused to our customers and our developer community.
February 2024(3 incidents)
Investigating new product purchasing
2 updates
Between 28th Feb 2024 23:15 UTC to 29th Feb 2024 00:05 UTC, we experienced issue with new product purchasing for all products. All new sign up products have been successfully provision and confirmed issue has been resolved and the service is operating normally.
We are investigating an issue with new product purchasing that is impacting for all products. Customers adding new cloud products may have experienced a long waiting page or an error page after attempting to add a product. We have mitigated the root cause and are working to resolve impact for customers who attempted to add a product during the impact period. We will provide more details within the next hour.
Service Disruptions Affecting Atlassian Products
5 updates
### **Summary** On February 14, 2024, between 20:05 UTC and 23:03 UTC, Atlassian customers on the following cloud products encountered a service disruption: Access, Atlas, Atlassian Analytics, Bitbucket, Compass, Confluence, Ecosystem apps, Jira Service Management, Jira Software, Jira Work Management, Jira Product Discovery, Opsgenie, StatusPage, and Trello. As part of a security and compliance uplift, we had scheduled the deletion of unused and legacy domain names used for internal service-to-service connections. Active domain names were incorrectly deleted during this event. This impacted all cloud customers across all regions. The issue was identified and resolved through the rollback of the faulty deployment to restore the domain names and Atlassian systems to a stable state. The time to resolution was two hours and 58 minutes. ### **IMPACT** External customers started reporting issues with Atlassian cloud products at 20:52 UTC. The impact of the failed change led to performance degradation or in some cases, complete service disruption. Symptoms experienced by end-users were unsuccessful page loads and/or failed interactions with our cloud products. ### **ROOT CAUSE** As part of a security and compliance uplift, we had scheduled the deletion of unused and legacy domain names that were being used for internal service-to-service connections. Active domain names were incorrectly deleted during this operation. ### **REMEDIAL ACTIONS PLAN & NEXT STEPS** We know that outages impact your productivity. The detection was delayed because existing testing & monitoring focused on service health rather than the entire system’s availability. To prevent a recurrence of this type of incident, we are implementing the following improvement measures: * Canary checks to monitor the entire system availability. * Faster rollback procedures for this type of service impact. * Stricter change control procedures for infrastructure modifications. * Migration of all DNS records to centralised management and stricter access controls on modification to DNS records. We apologize to customers whose services were impacted during this incident; we are taking immediate steps to improve the platform’s performance and availability. Thanks, Atlassian Customer Support
We experienced increased errors on Confluence, Jira Work Management, Jira Service Management, Jira Software, Opsgenie, Trello, Atlassian Bitbucket, Atlassian Access, Jira Align, Jira Product Discovery, Atlas, Compass, and Atlassian Analytics. The issue has been resolved and the services are operating normally.
We have identified the root cause of the Service Disruptions affecting all Atlassian products and have mitigated the problem. We are now monitoring this closely.
We have identified the root cause of the increased errors and have mitigated the problem. We continue to work on resolving the issue and monitoring this closely.
We are investigating reports of intermittent errors for all Cloud Customers across all Atlassian products. We will provide more details once we identify the root cause.
Trello is slow or unavailable
7 updates
### Summary On Feb. 13, 2024, between 8:00 AM and 11:34 AM UTC, Trello experienced severely degraded performance appearing as a full or partial outage to Atlassian customers. The event was triggered by a buildup of long-running queries against our database, leading to slowed API response times and causing Trello to be degraded or unavailable for users. The root cause of the incident was identified as a compression change in our database deployed approximately 11 hours earlier during a low-traffic period. As European customers came online, traffic started increasing, resulting in a buildup of queries and the subsequent incident. The incident was detected by our monitoring system at 8:07 AM UTC and was mitigated by reverting the compression change and restarting components of our database system. The total time to resolution was 3 hours and 34 minutes. ### **IMPACT** The overall impact was between 8:00 AM and 11:34 AM UTC on Feb. 13, 2024. The incident caused Trello to be fully or partially unavailable for customers using or attempting to access the site during this period. ### **ROOT CAUSE** The issue was caused by a compression change in our database, which resulted in the build up of queries in the system. This build up then caused API response times to increase to critical levels. During the incident many users received HTTP 429 errors as the system began rate-limiting in an attempt to recover. Users that did not receive errors experienced API response times 10-100x slower than our standard response times. ### **REMEDIAL ACTIONS PLAN & NEXT STEPS** We know that outages impact your productivity. We are prioritizing the following actions to avoid repeating this incident and reduce time to resolution: * Improve our process for releasing incremental configuration changes which would have allowed the team to identify the root cause before a peak load period and prevent similar incidents. * Adjust priority level of alerts related to this class of incident to improve signal to noise ratio and drive faster time to resolution. We apologize to customers whose services were impacted during this incident; we are taking immediate steps to improve Trello’s performance and availability. Thanks, Atlassian Customer Support
Trello is now available. Thank you for your patience.
We've identified an issue that was causing slowness in Trello and put a fix in place. We're seeing things improving for all our customers and we'll keep monitoring this latest fix for now. Once again, we appreciate your understanding!
Our team is still looking into the issue causing Trello to be slow or unavailable. Thank you for your understanding, we're working to have Trello back up and running as quickly as possible.
We're still investigating the root cause of this problem, and a new update will be shared soon! Thanks for your patience!
Our team is still investigating the issue that's causing Trello to be slow or unavailable, and working to bring Trello back up to speed as quickly as possible. Thanks for your patience and understanding!
We've noticed that Trello is responding slowly. This will be present in both the web and mobile apps. Our engineering team is actively investigating this incident and working to bring Trello back up to speed as quickly as possible. We'll keep you posted with further updates on this page.
December 2023(2 incidents)
Intermittent issues viewing and loading attachments
3 updates
This incident has been resolved.
A fix has been implement and we are monitoring the results.
Some users may see intermittent issues viewing and loading attachments. We are investigating the cause - for now, refreshing your browser may fix the issue. We will update this page with more information as it is discovered.
Egress connectivity timing out
5 updates
The systems are stable after the fix and monitoring for a specified duration
The issue was identified and a fix implemented. We are monitoring currently.
We are currently investigating an incident that result in outbound connections from Atlassian cloud in us-east-1 intermittently timing out. This affects Jira, Trello, Confluence, Ecosystem products. The features affected for these products are those that require opening a connection from Atlassian Cloud to public endpoints on the Internet
Including Atlassian Developer
We are currently investigating an incident that result in connection time outs on service egress proxy. This affects Jira, JSM, Confluence, BitBucket, Trello, Ecosystem products. The features affected for these products are those that require a connection to service egress.
November 2023(4 incidents)
Trello is unavailable
8 updates
### **SUMMARY** On Nov 30 2023, between 14:04 and 16:57 UTC, Atlassian customers using Trello experienced errors when accessing and interacting with the application. This incident impacted Trello users on the iOS and Android mobile apps as well as those using the Trello web app. The event was triggered by the release of a code change that eventually overloaded a critical part of the Trello database. The incident was detected immediately by our automated monitoring systems and was mitigated by disabling the relevant code change. The issue was extended by the failure of a secondary service whose recovery caused an increase in load on the same critical part of the Trello database, which created a negative feedback loop. This secondary service recovery involved reestablishing over a million connections, with each connection attempt adding load to the same part of the Trello database. We attempted to aid the service recovery by intentionally blocking some of the inbound Trello traffic to reduce load on the database and by increasing the capacity of the Trello database to better handle the high load. Over time the connections were all successfully reestablished, which returned Trello to a known good state. The total time to resolution was just under 3 hours. ### **IMPACT** The overall impact was between Nov 30 2023, 14:04 UTC and Nov 30 2023, 16:57 UTC on the Trello product. The incident caused service disruption to all Trello customers. Our metrics show there were elevated API response times and increased error rates through the entire incident period, which indicates that most users were unable to load Trello at all or easily interact with the application in any way. The particular database collection that was overloaded was one that is necessary for the Trello service to make authorization decisions, which meant that all requests were impacted. ### **ROOT CAUSE** The issue was caused by a series of changes intended to standardize Trello’s approach to authorizing requests, but had the unintended side effect of modifying a database query from a [targeted operation](https://www.mongodb.com/docs/manual/core/sharded-cluster-query-router/#std-label-sharding-mongos-targeted) to a [broadcast operation](https://www.mongodb.com/docs/manual/core/sharded-cluster-query-router/#std-label-sharding-mongos-broadcast). Broadcast operations are more resource-intensive as they must be sent to _all_ database servers to be satisfied. These broadcast operations eventually overloaded some of the Trello database servers as Trello approached its daily peak usage period on Nov 30 2023. 1. The first change of this type was deployed over a period of seven days at the end of August and changed the authorization type used by our websocket service. This meant that newly established websocket connections required this new broadcast query. At any given moment, we have a great deal of _established_ websocket connections, but the usual rate of _new_ websocket connections is relatively low. Therefore, our monitoring systems only detected a slight increase in resource usage and flagged this change as a low priority performance regression. We acknowledged the regression and created a task to identify and reduce the resource demands of these new queries. 2. The second change of this type was deployed over the course of a few days before being fully rolled out on Nov 29, 2023, the day before this incident. This change caused the Trello application server to use the new broadcast query while authorizing standard web browser traffic, which is the vast majority of our traffic. The change was fully deployed at 19:34 UTC on Nov 29, which was during a low traffic period. The next day, as the application approached its daily peak traffic period, our monitoring on the database servers indicated they were overloaded. When these database nodes were overloaded, users' HTTP requests received very slow responses or HTTP 504 errors. As we activated our load shedding strategies, some users received HTTP 429 errors. The incident’s length can be attributed to a secondary failure where our websocket servers experienced a rapid increase in memory leading to processes crashing with OutOfMemoryErrors. As new servers came online and the websockets attempted to reconnected, they once again generated the broadcast queries on the Trello database servers. These broadcast queries continued to put load on the database, which meant the Trello API continued to have high latency, thus perpetuating the negative feedback loop. We are working to determine the root cause of the OutOfMemoryErrors. We also determined after the incident that due to the Trello application server making the load shedding decision AFTER performing the authorization step, the overloaded database servers were still being queried before the request was rejected. We are working to improve our load shedding strategies post incident. ### **REMEDIAL ACTIONS PLAN & NEXT STEPS** We know that outages impact your productivity and we are continually working to improve our testing and preventative processes to prevent similar outages in the future. We are prioritizing the following improvement actions to avoid repeating this type of incident: * Increase the capacity of our database \(completed during the incident\). * This action is the most critical and is aimed at preventing a recurrence of this particular incident and gracefully recover if the websocket service were to fail again. * Refactor the new authorization approach to avoid [broadcast operations](https://www.mongodb.com/docs/manual/core/sharded-cluster-query-router/#std-label-sharding-mongos-broadcast). * Add pre-deployment tests to avoid releasing unnecessary broadcast operations. * Determine the root cause of the secondary failure of the websocket service. Furthermore, we deploy our changes only after thorough review and automated testing, and we deploy them progressively using feature flags to avoid broad impact. To minimize the impact of breaking changes to our environments, we will implement additional preventative measures: * Ensure that our load-shedding strategies fail fast. * Add monitoring to observe [broadcast operations](https://www.mongodb.com/docs/manual/core/sharded-cluster-query-router/#std-label-sharding-mongos-broadcast) in all our environments. We apologize to customers who were impacted during this incident; we are taking immediate steps to improve the platform’s performance and availability. Thanks, Atlassian Customer Support
The fix we released was successful, and all the issues our users were experiencing with Trello have been resolved. Thank you for your patience and understanding!
We've implemented a new fix to the issue affecting our database, and we're now seeing signs of recovery for all users on Trello. We'll keep monitoring this latest fix for now. Once again, we appreciate your understanding!
Our engineering team has identified an issue with Trello's database, and we're now working on implementing a fix to restore Trello's availability to all users. We appreciate everyone's understanding!
Our teams are still investigating the issue affecting Trello's availability, and a new update will be provided soon. We appreciate your understanding!
Our team is still investigating the issue that's affecting Trello's availability, and we'll provide a new update soon. Thanks for your patience and understanding!
We're still investigating the root cause of this problem, and a new update will be shared soon! Thanks for your patience!
We're currently investigating an issue that's causing Trello to be slow or unavailable to our users. A new update will be shared soon.
Forge Function Invocations outage impacting Smartlinks
1 update
Forge Invocations had an 8 minute outage between 2023-11-29 03:05:13 UTC to 2023-11-29 03:13:27 UTC resulting in Smart Links failing. This service has recovered post this time period.
Degraded experience for some of Trello's users
4 updates
This incident has been resolved. If you're still seeing issues, please reach out at https://trello.com/contact/
Our Engineers have rolled out a fix and have seen services recovering. We're now monitoring it.
Our Engineers have identified the issue and are working on a fix. We have found the following services to be affected: - Inability to create new Automation commands - Card changes not saving - Some file uploads failing
We are currently investigating an issue affecting some of Trello's services. We have identified that some users may experience trouble with the following service: - Automations commands not running - Card changes not saving
Trello is down
5 updates
We want to let you know that the changes we made on our side fixed all the connection issues our users were experiencing with Trello, and all services are functional right now. We appreciate your patience and understanding.
The changes we released have fixed the connection issues our users were experiencing when connecting to Trello. We're now monitoring it.
We've identified the root cause of the issue, and we made some adjustments, which should bring Trello's functionalities back soon.
We are continuing to work on a fix for this issue.
We've identified an issue on Trello that's affecting multiple users to connect to our servers. We'll provide a new update soon.
📡 Tired of checking Trello API status manually?
Better Stack monitors uptime every 30 seconds and alerts you instantly when Trello API goes down.