fastmail.xml (387721B)
1 <?xml version='1.0' encoding='utf-8'?> 2 <feed xmlns='http://www.w3.org/2005/Atom' xml:base='https://www.fastmail.com'> 3 <title>Fastmail Blog</title> 4 <subtitle>Blog posts from the Fastmail team</subtitle> 5 <updated>2025-03-21T00:00:01Z</updated> 6 <id>https://www.fastmail.com</id> 7 <link rel='alternate' type='text/html' hreflang='en' href='https://www.fastmail.com/blog/' /> 8 <link rel='self' type='application/atom+xml' href='https://www.fastmail.com/blog/feed.xml' /> 9 <rights>© 2025 Fastmail Pty. Ltd. All rights reserved</rights><entry> 10 <title>Not OK, Cupid</title> 11 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/not-ok-cupid/' /> 12 <id>https://www.fastmail.com/blog/not-ok-cupid/</id> 13 <updated>2025-03-21T00:00:01Z</updated><author> 14 <name>Bron Gondwana</name> 15 </author><content xml:lang='en' type='html'><p>I don’t usually like to call out the bad behaviour of specific companies, but the egregious mis-design and lack of acknowledging it justify this case.</p><h2 id="welcome-to-ok-cupid" tabindex="-1">Welcome to OkCupid</h2><p>A couple of weeks ago, I started seeing many “Welcome to OkCupid” emails, both on my personal address and a couple of related addresses, but also to multiple Fastmail official contact addresses — legal, partnerships, press, etc. Specifically, this list included <code>trash@brong.net</code> — an address that has never been used to send or receive email and appears in precisely one place — <a href="https://www.fastmail.com/blog/a-tangled-path-of-workarounds/" target="_blank" rel="noopener">an article on our blog</a>! It seems quite clear that somebody scraped our website and used the addresses to sign up. I’m aware of at least 10 addresses, but there are likely others that either go to someone else or addresses that no longer exist.</p><p>It didn’t stop there, though. I’ve been getting tons of “someone likes you”, “you have an intro,” and even an “IMPORTANT: We removed your photo on OkCupid.” email saying that inappropriate content was posted to “our” account!</p><h2 id="the-real-world-consequences-of-poor-email-validation" tabindex="-1">The real-world consequences of poor email validation</h2><p>This isn’t just an inconvenience — it has real security implications. Websites that fail to properly validate email ownership can be exploited for malicious purposes. Attackers can use unverified sign-ups to flood inboxes, making it easier to hide critical emails among the noise — something we’ve discussed our own experience of in our post on <a href="https://www.fastmail.com/blog/when-two-factor-authentication-is-not-enough/" target="_blank" rel="noopener">2FA vulnerabilities</a>. There are established <a href="https://www.m3aawg.org/sites/default/files/document/M3AAWG_Senders_BCP_Ver3-2015-02.pdf" target="_blank" rel="noopener">best practices</a> (PDF) for handling email sign-ups responsibly, practices that OkCupid is failing to follow.</p><h2 id="no-way-out" tabindex="-1">No way out</h2><p>When I tried to unsubscribe using the one-click unsubscribe button in one of the emails, I was met with an error: “Something went wrong, please try again later.”</p><p>Curious, I tried to recover a password on one of these accounts (the one with my personal email address) and successfully changed the password. Then, I was asked to confirm my login with a message sent to the number associated with the account. A number I didn’t know. A number that wasn’t mentioned on that page, so I still don’t know anything about it — not even which country it was from.</p><p>This raises further security concerns; the attacker could have also caused random recovery numbers to be texted to another poor victim’s phone. Alternatively, they could confirm that my email address is actively monitored, increasing its value for further attacks. Either way, what I couldn’t do was actually close the account.</p><h2 id="whack-a-mole" tabindex="-1">Whack-a-mole</h2><p>So, I contacted OkCupid’s support. Here’s what they said:</p><p><em>I’ve removed the user from the site and banned the email address to prevent any new accounts from being created. That should resolve the issue, but if you encounter anything like this again in the future, please don’t hesitate to reach out, and we’ll address it right away.</em></p><p>So, I need to contact support manually for each new email address. This is neither scalable nor acceptable; people don’t have this amount of time.</p><p>Furthermore, my email address is now on another random blocklist somewhere on the internet, where I have no control and no way to unblock it. I don’t anticipate wanting to use OkCupid’s service, but if I did in the future, I would have to go through another dance to get the address unlocked again — or more likely, treat that particular email address as soiled and create another one.</p><h2 id="not-ok" tabindex="-1">Not OK</h2><p>So I say, not OK, OkCupid. Not OK.</p><p>The usefulness of email depends on responsible behaviour from all service providers. Companies that engage in shady or outright inappropriate practices make the internet worse for everyone.</p><p>OkCupid’s failure to implement even the <a href="https://en.wikipedia.org/wiki/Opt-in_email#Confirmed_opt-in_(COI)/double_opt-in_(DOI)" target="_blank" rel="noopener">simplest form</a> of email validation is unacceptable. Until they address these issues properly (not through the support response provided here), they remain part of the problem, not the solution.</p><h2 id="could-we-have-avoided-this" tabindex="-1">Could we have avoided this?</h2><p>In this case, we published those addresses online. There’s always a risk of receiving spam when you do that, one could even reasonably say “we were asking for it”. We expected spam. If you want to reduce your risk of being spammed, it helps to not publish your email address on the public web!</p><p>What we we didn’t was expect a relatively reputable service being used to facilitate us being spammed.</p><p>One great protection is using different address for each different organisation you deal with — that way if your address leaks (or they sell it), you know where the breach happened, and you can more easily block just the problem messages.</p><p>Fastmail’s masked email feature is a great way to implement this strategy. Masked emails are designed, particularly when integrated with a password manager, to make it very easy to create new addresses, and track where they are expected to be used.</p><p>Being a good internet citizen is one of <a href="https://www.fastmail.com/company/values/" target="_blank" rel="noopener">Fastmail’s core values</a>. We require verification for sending identities, ensuring that only legitimate users can send from an address they claim they own. This is the level of responsibility every email provider should uphold, and we applaud the others who also do.</p></content> 16 </entry><entry> 17 <title>The evolution of the advanced fee scam</title> 18 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/the-evolution-of-the-advanced-fee-scam/' /> 19 <id>https://www.fastmail.com/blog/the-evolution-of-the-advanced-fee-scam/</id> 20 <updated>2025-02-14T05:00:00Z</updated><author> 21 <name>Aric Archebelle-Smith</name> 22 </author><content xml:lang='en' type='html'><p>As one of Fastmail’s customer support agents, part of my job is making sure that our customers are well-informed about rising trends in fraud so that they can be sure to steer clear of them. While our customers tend to be tech-savvy enough to spot the average scam email from a mile away, online scammers grow increasingly more sophisticated every year.</p><p>I recently attended the 62nd General Meeting of the <a href="https://www.m3aawg.org/" target="_blank" rel="noopener">Messaging, Malware, Mobile Anti-Abuse Working Group (M3AAWG)</a> in Toronto. There, I spoke with others in the email and anti-abuse industry about the increase in advanced fee scams they’d observed in the years since the onset of the COVID-19 pandemic.</p><p>Advanced fee scams are not a new type of scam, but scammers have begun running a much more sophisticated version of this old-school scam. One that can convince even those who know to be cautious when navigating the internet.</p><p>Historically, advanced fee scams involved scammers promising the victim some sort of too-good-to-be-true opportunity or reward. The only catch is that the the victim has to pay a fee before they can receive the promised reward or opportunity. Generally, the scammer claims this fee is just to cover processing fees, background checks, training materials, or some other reasonable sounding expense. They assure the victim that they’ll be reimbursed for this expense down the road. Once the victim pays the fee, the scammer goes silent and the victim realizes that they’ve been conned.</p><p>Until recently, advanced fee scams were your garden variety “Nigerian prince” scam that savvy internet users quickly learned to avoid. Someone would offer the victim a large payoff if the victim could just cover the relatively small wire transfer or bank processing fees. For most email users, this type of con was easy to detect and most people knew to watch out for them.</p><p>Advanced fee scams have recently evolved to masquerade as hiring and work-from-home opportunities, targeting people who are looking for work in an already highly competitive job market. The scammers will pose as hiring managers or recruiters, and will even go so far as to reach out to victims over legitimate hiring websites, such as LinkedIn.</p><p>The victim is led to believe that they are being considered for a job or internship opportunity, but they’ll be asked to pay a fee as part of the hiring process. In some cases, the victim is given a link to the company’s preferred online vendor, where they are told to purchase the items they’ll need for the job. The scammer tells the victim that they’ll be reimbursed for these purchases later. However, the link takes the victim to a fake webstore where the payment is taken, but no goods are ever sent. At this point, the scammer stops responding to the victim.</p><p>More frequently, the scammers ask the victim to pay a small fee to cover some other aspect of the hiring process. Generally, the scammer will claim this is an application fee or something similar. Of course, the scammer stops responding to the victim’s messages as soon as they receive the payment.</p><p>In some cases, scammers will even conduct actual phone or video interviews with the victim as part of the phony hiring process. There’s no way to know how this data is being used by the attackers without insider knowledge.</p><p>This combination of fraudulent hiring and advanced fee scams allows attackers to collect both money and personally identifying information from vulnerable populations.</p><p>I recommend the following precautions to avoid becoming the victim of one of these scams:</p><ul> <li>Confirm the legitimacy of any jobs you are interested in applying for by verifying that the position is listed on the company’s website.</li> <li>Make sure that any emails you receive from a hiring manager are actually coming from the company’s domain or from the domain of a legitimate staffing agency. Double check that there are no typos or <a href="https://itservices.wp.st-andrews.ac.uk/2024/03/12/identifying-fraudsters-using-cyrillic-characters" target="_blank" rel="noopener">look-alike characters</a> in the domain.</li> <li>Take the same precautions with any URLs that are shared with you via email or on hiring sites. Scammers can set up convincing look-alike websites, but you can check the URL to verify that you are being directed to the company’s legitimate website.</li> <li>Even if a message appears to be sent from a company’s actual domain, there’s a chance that the message could be spoofed, meaning the scammer forged the email’s “From” address to make it look like it came from a certain person or company. Chances are that these messages would get flagged as spam, but it’s still a good idea to confirm that a message hasn’t been spoofed by checking the headers of the message. Fastmail makes it easy to view the full headers of a message. Simply click the <strong>Actions</strong> drop down and select <strong>Show raw message</strong> to see the full headers of the message and verify that the message passed <a href="https://www.fastmail.help/hc/en-us/articles/1500000280461-Sender-authentication" target="_blank" rel="noopener">sender authentication checks</a>. If you’re not familiar with how to read email headers, you can always reach out to Fastmail’s friendly and knowledgeable <a href="https://support.fastmail.com/support/" target="_blank" rel="noopener">support team</a> to help confirm a message’s legitimacy.</li> <li>If a job opportunity seems too good to be true, or you’re told that you’ve been accepted for a position almost immediately with little to no interview process, chances are the hiring manager or recruiter that you’re talking with is actually a scammer.</li> <li>If at any point in the interview process the recruiter asks to stop communicating via email and asks you to contact them on Telegram, WhatsApp, or any other end-to-end encrypted communication platform, they are almost certainly trying to scam you.</li> <li>If the company requires payment from you for a job opportunity, we ultimately recommend that you do not proceed. It’s extraordinarily rare for a legitimate company to require payment from you for a job opportunity.</li> </ul><p>As these scams become more pervasive, it’s crucial that those on the job market educate themselves on the potential scams that are out there. Knowing how to recognize and avoid these fraudulent job listings can ensure you don’t waste your time, lose money, or divulge your personal data to scammers.</p></content> 23 </entry><entry> 24 <title>Dec 24: Twenty five years of Fastmail</title> 25 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/twenty-five-years-of-fastmail/' /> 26 <id>https://www.fastmail.com/blog/twenty-five-years-of-fastmail/</id> 27 <updated>2024-12-24T00:00:01Z</updated><author> 28 <name>Rob Mueller</name> 29 </author><author> 30 <name>Bron Gondwana</name> 31 </author><content xml:lang='en' type='html'><p>This is the twenty-fourth and final post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/ten-years-of-jmap/">Dec 23: Ten years of JMAP</a>. Thanks for reading, see you again next year.</p><p>As we conclude this year’s Advent posts, we are reflecting back over 25 years! Fastmail was founded in 1999, to fill a gap which existed at the time — in the space between ISPs, slow and ad-riddled free email services, and clunky, bloated Enterprise systems, there was no professional email service for a small business or sophisticated email user.</p><p>So we built one! Fastmail: a slick, professional, web-based email service.</p><p>In the 25 years since, we have seen many changes in the email landscape and the world around us. The advent of Gmail and conversations as a standard email model. The rise of encryption focused services like Protonmail (with the <a href="https://www.fastmail.com/features/security/" target="_blank" rel="noopener">pros and cons</a> of storing email as opaque, unsearchable blobs). Highly opinionated “reinventions” of email like Hey. And of course the multiple premature announcements that email was dead, to be replaced by the latest new craze.</p><p>We were purchased by Opera Software in 2010, but after some changes in Opera’s strategic direction, thankfully a handful of the staff managed to buy the company back in 2013. We then purchased another email service Pobox in 2015, who had been running an email service even longer than us. We have recently finished merging their product into our system; who knew it was going to take so long to integrate everything!</p><p>Through all of this we’ve been grateful to have such loyal customers. We regularly hear from customers how much they appreciate the Fastmail service. Our fantastic customer support. The continuous, thoughtful, and well designed improvements to our product. The high performance and reliability of our service. The ongoing commitment to integrity, privacy, and longevity.</p><p>The result is that we have a greater than 90% annual renewal rate, and an ongoing stream of new customers from the word of mouth recommendations of existing customers. We have and continue to grow every year in a sustainable and deliberate way. We’re insanely grateful for this. We get to focus on making email better for our customers, to work with and build cool technology — with really smart colleagues. We can solve complex problems, build well-designed solutions, and improve email standards without having to always hustle for the next sale.</p><p>It’s an enviable position to be in. Email remains the largest open federated communication network on the internet. Not controlled by a single company. Not part of any walled garden that can change at any time. Through open standards, email allows you to choose the best provider and to move your email where is best for you. As we said in our first post of this series, we will continue to “Make email better”, for our customers and for everyone.</p><p>We love our work, and the customers who trust us with their email and make this all possible. So cheers to you, Fastmail’s customers. We get to make email better, the product you use and the ecosystem we all operate in, while having fun and working on interesting problems with great people.</p><p>Here’s to another 25 years.</p></content> 32 </entry><entry> 33 <title>Dec 23: Ten years of JMAP</title> 34 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/ten-years-of-jmap/' /> 35 <id>https://www.fastmail.com/blog/ten-years-of-jmap/</id> 36 <updated>2024-12-23T00:00:01Z</updated><author> 37 <name>Bron Gondwana</name> 38 </author><content xml:lang='en' type='html'><p>This is the twenty-third post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/why-we-use-our-own-hardware/">Dec 22: Why we use our own hardware at Fastmail</a>. The final post is <a href="/blog/twenty-five-years-of-fastmail/">Dec 24: Twenty five years of Fastmail</a>.</p><p>Exactly 10 years ago, we <a href="/blog/jmap-a-better-way-to-email/">announced JMAP on our blog</a>, along with a <a href="https://youtu.be/8qCSK-aGSBA" target="_blank" rel="noopener">video by baby-faced Bron and Neil</a>!</p><p>JMAP: A better way to email. We knew it would be a long road, but we’re really glad we did it and created an open standard rather than staying with our own custom protocol.</p><h2 id="some-moments-along-the-way" tabindex="-1">Some moments along the way</h2><p>We started by workshopping the idea around the industry. I did a <a href="https://www.youtube.com/watch?v=yyXlUR1hbr4" target="_blank" rel="noopener">lightning talk at OSCON in 2014</a>, our first attempt to find developers who could give us feedback on our design. By far the best find was <a href="https://rjbs.cloud/" target="_blank" rel="noopener">Ricardo Signes</a>, Pobox developer, who I met at a bar on the last day! This led to us acquiring the product and (our main goal) acqui-hiring Rik, who is now one of the company owners, as well as a JMAP enthusiast!</p><p><picture><source type="image/webp" srcset="/assets/images/oscon-bron-rik-91PNJqGmRJ-375.webp 375w, /assets/images/oscon-bron-rik-91PNJqGmRJ-750.webp 750w, /assets/images/oscon-bron-rik-91PNJqGmRJ-1154.webp 1154w" sizes="(max-width: 425px) 375px, 750px"><img alt="Bron and Rik at OSCON in 2014" loading="lazy" decoding="async" src="/assets/images/oscon-bron-rik-91PNJqGmRJ-375.png" width="1154" height="893" srcset="/assets/images/oscon-bron-rik-91PNJqGmRJ-375.png 375w, /assets/images/oscon-bron-rik-91PNJqGmRJ-750.png 750w, /assets/images/oscon-bron-rik-91PNJqGmRJ-1154.png 1154w" sizes="(max-width: 425px) 375px, 750px"></picture></p><p>Neil and I attended <a href="https://web.archive.org/web/20150908015219/http://inboxlove.com/" target="_blank" rel="noopener">Inbox Love</a> in the Bay Area in 2014 as well. This gave us a chance to meet some of our technical peers in the big companies, relationships which we have continued to foster over the years. This hasn’t led to everyone dropping everything and implementing our protocols, but it has led to some collaborative design and ongoing conversations, and I believe its has prevented a proliferation of other protocols since people point to JMAP instead of inventing a new thing themselves. We also were told “go to the IETF”, but the IETF seemed big and scary and we didn’t know how, so that took a while.</p><p>Instead, we joined <a href="https://www.calconnect.org/" target="_blank" rel="noopener">CalConnect</a> and started working on Calendar formats and standards, while promoting JMAP more generally. Eventually we made more contacts in the IETF, and finally in 2017 went to our first IETF meeting in Chicago. At this point, the <a href="https://datatracker.ietf.org/wg/jmap/history/" target="_blank" rel="noopener">JMAP working group</a> was born.</p><p>In the crucible of the IETF, we made major changes. The authentication was removed. Method names were split into <code>Object/action</code> and a ton of smaller changes were made. The <a href="https://www.rfc-editor.org/rfc/rfc8620.html" target="_blank" rel="noopener">Core</a> and <a href="https://www.rfc-editor.org/rfc/rfc8621.html" target="_blank" rel="noopener">Mail</a> JMAP specifications were published in 2019, and then we got to work on the rest of the stack.</p><p>JMAP <a href="https://www.rfc-editor.org/rfc/rfc9610.html" target="_blank" rel="noopener">Contacts</a> was published just last week, and JMAP <a href="https://datatracker.ietf.org/doc/draft-ietf-jmap-calendars/" target="_blank" rel="noopener">Calendars</a> is very close to being published. I’m also keen to add <a href="https://datatracker.ietf.org/doc/draft-ietf-jmap-filenode/" target="_blank" rel="noopener">Filenode</a> support, but we want to get more experience with other filesystem providers before we standardize that (it’s currently based very closely on Fastmail’s own custom Node objects for our filestorage feature).</p><h2 id="what-s-next" tabindex="-1">What’s next</h2><p>We created JMAP because we could see that without it, the email world was going to become more insular, with the only modern standards for email access being proprietary. With Calendars and Contacts, we’re bringing the same easy-to-use JSON objects under a single protocol.</p><p>We started the <a href="https://makebetter.email/" target="_blank" rel="noopener">Make Better Email</a> conference last year, focused on improving the authentication workflow and also promoting JMAP usage. It’s a very small, invite-only conference where we do deep technical design work on improving interoperability and discoverability between clients and services. It was in Philadelphia last year, London this year, and we expect to be in Philadelphia again next year — likely in mid November after <a href="https://www.ietf.org/meeting/124/" target="_blank" rel="noopener">IETF 124</a> so we don’t cross over Halloween. If you think you’d be a useful addition to the meeting, pop us an email via the link at the bottom of the site.</p><p>Some work products of the previous conferences have been:</p><ul> <li><a href="https://datatracker.ietf.org/doc/draft-jenkins-oauth-public/" target="_blank" rel="noopener">An OAuth profile for open-protocol clients</a></li> <li><a href="https://datatracker.ietf.org/doc/draft-jenkins-emailpush/" target="_blank" rel="noopener">A profile for JMAP push</a> allowing you get some of the benefits of JMAP’s push capability without having to do a full JMAP implementation</li> <li><a href="https://datatracker.ietf.org/doc/draft-ietf-mailmaint-autoconfig/" target="_blank" rel="noopener">A specification for auto-discovery of configuration information</a></li> </ul><p>Over the past year some of us have also been working in the server-to-server space with <a href="https://datatracker.ietf.org/doc/draft-gondwana-dkim2-motivation/" target="_blank" rel="noopener">an idea</a> that may wind up replacing or enhancing DKIM.</p><p>And finally, next year we will be investing a lot more effort into making the <a href="https://www.cyrusimap.org/" target="_blank" rel="noopener">Cyrus IMAP</a> server not just a reference implementation for JMAP, but much easier to both develop and run.</p><p>In 10 years time, I hope to post about how Cyrus and JMAP have taken over the world, but I’ll also happily settle for them having both improved Fastmail’s product immeasurably, having plenty of happy customers, and continuing to help make email better for everybody through our work.</p></content> 39 </entry><entry> 40 <title>Dec 22: Why we use our own hardware at Fastmail</title> 41 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/why-we-use-our-own-hardware/' /> 42 <id>https://www.fastmail.com/blog/why-we-use-our-own-hardware/</id> 43 <updated>2024-12-22T00:00:01Z</updated><author> 44 <name>Rob Mueller</name> 45 </author><content xml:lang='en' type='html'><p>This is the twenty-second post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/fastmail-in-a-box/">Dec 21: Fastmail in a box</a>. The next post is <a href="/blog/ten-years-of-jmap/">Dec 23: Ten years of JMAP</a>.</p><h2 id="why-we-use-our-own-hardware" tabindex="-1">Why we use our own hardware</h2><p>There has recently been talk of <a href="https://www.google.com/search?q=Cloud+Repatriation" target="_blank" rel="noopener">cloud repatriation</a> where companies are moving from the cloud to on premises, with some particularly <a href="https://basecamp.com/cloud-exit" target="_blank" rel="noopener">noisy examples</a>.</p><p>Fastmail has a long history of using our <a href="https://www.fastmail.com/blog/standalone-mail-servers/" target="_blank" rel="noopener">own</a> <a href="https://www.fastmail.com/blog/getting-the-most-out-of-hardware/" target="_blank" rel="noopener">hardware</a>. We have over two decades of experience running and optimising our systems to use our own <a href="https://en.wikipedia.org/wiki/Bare-metal_server" target="_blank" rel="noopener">bare metal</a> servers efficiently.</p><p>We get way better cost optimisation compared to moving everything to the cloud because:</p><ol> <li>We understand our short, medium and long term usage patterns, requirements and growth very well. This means we can plan our hardware purchases ahead of time and don’t need the fast dynamic scaling that cloud provides.</li> <li>We have in house operations experience installing, configuring and running our own hardware and networking. These are skills we’ve had to maintain and grow in house since we’ve been doing this for 25 years.</li> <li>We are able to use our hardware for long periods. We find our hardware can provide useful life for anywhere from 5-10 years depending on what it is and when in the global technology cycle it was bought, meaning we can amortise and depreciate the cost of any hardware over many years.</li> </ol><p>Yes, that means we have to do more ourselves, including planning, choosing, buying, installing, etc, but the tradeoff for us has and we believe continues to be significantly worth it.</p><h2 id="hardware-over-the-years" tabindex="-1">Hardware over the years</h2><p>Of course over the 25 years we’ve been running Fastmail we’ve been through a number of hardware changes. For many years, our IMAP server storage platform was a combination of <a href="https://www.urbandictionary.com/define.php?term=Spinning%20Rust" target="_blank" rel="noopener">spinning rust</a> drives and <a href="https://www.areca.com.tw/" target="_blank" rel="noopener">ARECA RAID controllers</a>. We tended to use faster 15k RPM SAS drives in <a href="https://en.wikipedia.org/wiki/Standard_RAID_levels#RAID_1" target="_blank" rel="noopener">RAID1</a> for our hot meta data, and 7.2k RPM SATA drives in <a href="https://en.wikipedia.org/wiki/Standard_RAID_levels#RAID_6" target="_blank" rel="noopener">RAID6</a> for our main email blob data.</p><p>In fact it was slightly more complex than this. Email blobs were written to the fast RAID1 SAS volumes on delivery, but then a separate archiving process would move them to the SATA volumes at low server activity times. Support for all of this had been added into <a href="https://www.cyrusimap.org" target="_blank" rel="noopener">cyrus</a> and our tooling over the years in the form of separate “meta”, “data” and <a href="https://www.cyrusimap.org/3.8/imap/reference/admin/locations/archive-partitions.html" target="_blank" rel="noopener">“archive”</a> partitions.</p><h2 id="moving-to-nv-me-ssds" tabindex="-1">Moving to NVMe SSDs</h2><p>A few years ago however we made our biggest hardware upgrade ever. We moved all our email servers to a new <a href="https://www.supermicro.com/en/aplus/system/2u/2113/as-2113s-wn24rt.cfm" target="_blank" rel="noopener">2U AMD platform</a> with pure <a href="https://www.solidigm.com/products/data-center.html" target="_blank" rel="noopener">NVMe SSDs</a>. The density increase (24 x 2.5&quot; NVMe drives vs 12 x 3.5&quot; SATA drives per 2U) and performance increase was enormous. We found that these new servers performed even better than our initial expectations.</p><p>At the time we upgraded however NVMe RAID controllers weren’t widely available. So we had to decide on how to handle redundancy. We considered a RAID-less setup using raw SSDs drives on each machine with synchronous application level replication to other machines, but the software changes required were going to be more complex than expected.</p><p>We were looking at using classic Linux <a href="https://en.wikipedia.org/wiki/Mdadm" target="_blank" rel="noopener">mdadm RAID</a>, but the <a href="https://en.wikipedia.org/wiki/RAID#Atomicity" target="_blank" rel="noopener">write hole</a> was a concern and the <a href="https://docs.kernel.org/driver-api/md/raid5-cache.html" target="_blank" rel="noopener">write cache</a> didn’t seem well tested at the time.</p><p>We decided to have a look at <a href="https://arstechnica.com/information-technology/2020/05/zfs-101-understanding-zfs-storage-and-performance/" target="_blank" rel="noopener">ZFS</a> and at least test it out.</p><p>Despite some of the cyrus on disk database structures being fairly hostile to <a href="https://en.wikipedia.org/wiki/ZFS#Copy-on-write_transactional_model" target="_blank" rel="noopener">ZFS Copy-on-write</a> semantics, they were still incredibly fast at all the IO we threw at them. And there were some other wins as well.</p><h2 id="zfs-compression-and-tuning" tabindex="-1">ZFS compression and tuning</h2><p>When we rolled out ZFS for our email servers we also enabled <a href="https://freebsdfoundation.org/wp-content/uploads/2021/05/Zstandard-Compression-in-OpenZFS.pdf" target="_blank" rel="noopener">transparent Zstandard compression</a>. This has worked very well for us, saving about 40% space on all our email data.</p><p>We’ve also recently done some additional calculations to see if we could tune some of the parameters better. We sampled 1 million emails at random and calculated how many blocks would be required to store those emails uncompressed, and then with <a href="https://klarasystems.com/articles/tuning-recordsize-in-openzfs/" target="_blank" rel="noopener">ZFS record sizes</a> of 32k, 128k or 512k and zstd-3 or zstd-9 compression options. Although ZFS <a href="https://en.wikipedia.org/wiki/ZFS#ZFS's_approach:_RAID-Z_and_mirroring" target="_blank" rel="noopener">RAIDz2</a> seems conceptually similar to classic RAID6, the way it <a href="https://ibug.io/blog/2023/10/zfs-block-size/" target="_blank" rel="noopener">actually stores blocks of data</a> is quite different and so you have to take into account volblocksize, how files are split into logical recordsize blocks, and number of drives when doing calculations.</p><pre><code> Emails: 1,026,000 46 Raw blocks: 34,140,142 47 32k &amp; zstd-3, blocks: 23,004,447 = 32.6% saving 48 32k &amp; zstd-9, blocks: 22,721,178 = 33.4% saving 49 128k &amp; zstd-3, blocks: 20,512,759 = 39.9% saving 50 128k &amp; zstd-9, blocks: 20,261,445 = 40.7% saving 51 512k &amp; zstd-3, blocks: 19,917,418 = 41.7% saving 52 512k &amp; zstd-9, blocks: 19,666,970 = 42.4% saving 53 </code></pre><p>This showed that the defaults of 128k record size and zstd-3 were already pretty good. Moving to a record size of 512k improved compression over 128k by a bit over 4%. Given all meta data is cached separately, this seems a worthwhile improvement with no significant downside. Moving to zstd-9 improved compression over zstd-3 by about 2%. Given the CPU cost of compression at zstd-9 is about 4x zstd-3, even though emails are immutable and tend to be kept for a long time, we’ve decided not to implement this change.</p><h2 id="zfs-encryption" tabindex="-1">ZFS encryption</h2><p>We always enable <a href="https://en.wikipedia.org/wiki/Data_at_rest#Encryption" target="_blank" rel="noopener">encryption at rest</a> on all of our drives. This was usually done with <a href="https://en.wikipedia.org/wiki/Linux_Unified_Key_Setup" target="_blank" rel="noopener">LUKS</a>. But with ZFS this was <a href="https://arstechnica.com/gadgets/2021/06/a-quick-start-guide-to-openzfs-native-encryption/" target="_blank" rel="noopener">built in</a>. Again, this reduces overall system complexity.</p><h2 id="going-all-in-on-zfs" tabindex="-1">Going all in on ZFS</h2><p>So after the success of our initial testing, we decided to go all in on ZFS for all our large data storage needs. We’ve now been using ZFS for all our email servers for over 3 years and have been very happy with it. We’ve also moved over all our database, log and backup servers to using ZFS on NVMe SSDs as well with equally good results.</p><h2 id="ssd-lifetimes" tabindex="-1">SSD lifetimes</h2><p>The flash memory in SSDs has a finite life and <a href="https://en.wikipedia.org/wiki/Flash_memory#Write_endurance" target="_blank" rel="noopener">finite number of times it can be written to</a>. SSDs employ increasingly complex <a href="https://en.wikipedia.org/wiki/Wear_leveling" target="_blank" rel="noopener">wear levelling</a> algorithms to spread out writes and increase drive lifetime. You’ll often see the quoted endurance of an enterprise SSD as either an absolute figure of “Lifetime Writes”/“Total bytes written” like 65 PBW (petabytes written) or a relative per-day figure of “Drive writes per day” like 0.3, which you can convert to lifetime figure by multiplying by the drive size and the drive expected lifetime which is often assumed to be 5 years.</p><p>Although we could calculate IO rates for existing <a href="https://en.wikipedia.org/wiki/Hard_disk_drive" target="_blank" rel="noopener">HDD</a> systems, we were making a significant number of changes moving to the new systems. Switching to a COW filesystem like ZFS, removing the special casing meta/data/archive partitions, and the massive latency reduction and performance improvements mean that things that might have taken extra time previously and ended up batching IO together, are now so fast it actually causes additional separated IO actions.</p><p>So one big unknown question we had was how fast would the SSDs wear in our actual production environment? After several years, we now have some clear data. From one server at random but this is fairly consistent across the fleet of our oldest servers:</p><pre><code># smartctl -a /dev/nvme14 54 ... 55 Percentage Used: 4% 56 </code></pre><p>At this rate, we’ll replace these drives due to increased drive sizes, or entirely new physical drive formats (such <a href="https://www.snia.org/forums/cmsi/knowledge/formfactors" target="_blank" rel="noopener">E3.S</a> which appears to finally be gaining traction) long before they get close to their rated write capacity.</p><p>We’ve also anecdotally found SSDs just to be much more reliable compared to HDDs for us. Although we’ve only ever used <a href="https://www.micron.com/products/storage/ssd/data-center-ssd/" target="_blank" rel="noopener">datacenter</a> <a href="https://www.solidigm.com/products/data-center.html" target="_blank" rel="noopener">class</a> SSDs and <a href="https://www.seagate.com/www-content/datasheets/pdfs/exos-7-e8-data-sheet-DS1957-1-1709US-en_US.pdf" target="_blank" rel="noopener">HDDs</a> failures and replacements every few weeks were a regular occurrence on the old fleet of servers. Over the last 3+ years, we’ve only seen a couple of SSD failures in total across the entire upgraded fleet of servers. This is easily less than one tenth the failure rate we used to have with HDDs.</p><h2 id="storage-cost-calculation" tabindex="-1">Storage cost calculation</h2><p>After converting all our email storage to NVMe SSDs, we were recently looking at our data backup solution. At the time it consisted of a number of older 2U servers with 12 x 3.5&quot; SATA drive bays and we decided to do some cost calculations on:</p><ol> <li>Move to cloud storage.</li> <li>Upgrade the HD drives in existing servers.</li> <li>Upgrade to SSD NVMe machines.</li> </ol><h3 id="1-cloud-storage" tabindex="-1">1. Cloud storage:</h3><p>Looking at various providers, the per TB per month price, and then a yearly price for 1000Tb/1Pb (prices as at Dec 2024)</p><ul> <li><a href="https://aws.amazon.com/s3/pricing/" target="_blank" rel="noopener">Amazon S3</a> - $21 -&gt; $252,000/y</li> <li><a href="https://developers.cloudflare.com/r2/pricing/" target="_blank" rel="noopener">Cloudflare R2</a> - $15 -&gt; $180,000/y</li> <li><a href="https://wasabi.com/pricing" target="_blank" rel="noopener">Wasabi</a> - $6.99 -&gt; $83,880/y</li> <li><a href="https://www.backblaze.com/cloud-storage/pricing" target="_blank" rel="noopener">Backblaze B2</a> - $6 -&gt; $72,000/y</li> <li><a href="https://aws.amazon.com/s3/pricing/" target="_blank" rel="noopener">Amazon S3 Glacier Instant Retrieval</a> - $4 -&gt; $48,000/y</li> <li><a href="https://aws.amazon.com/s3/pricing/" target="_blank" rel="noopener">Amazon S3 Glacier Deep Archive (12 hour retrieval time)</a> - $0.99 -&gt; $11,880/y</li> </ul><p>Some of these (e.g. Amazon) have potentially significant bandwidth fees as well.</p><p>It’s interesting seeing the spread of prices here. Some also have a bunch of weird edge cases as well. e.g. “The S3 Glacier Flexible Retrieval and S3 Glacier Deep Archive storage classes require an additional 32 KB of data per object”. Given the large retrieval time and extra overhead per-object, you’d probably want to store small incremental backups in regular S3, then when you’ve gathered enough, build a biggish object to push down to Glacier. This adds implementation complexity.</p><ul> <li><em>Pros</em>: No limit to amount we store. Assuming we use S3 compatible API, can choose between multiple providers.</li> <li><em>Cons</em>: Implementation cost of converting existing backup system that assumes local POSIX files to S3 style object API is uncertain and possibly significant. Lowest cost options require extra careful consideration around implementation details and special limitations. Ongoing monthly cost that will only increase as amount of data we store increases. Uncertain if prices will go down or not, or even go up. Possible significant bandwidth costs depending on provider.</li> </ul><h3 id="2-upgrade-hdds" tabindex="-1">2. Upgrade HDDs</h3><p><a href="https://www.seagate.com/au/en/products/enterprise-drives/exos-x/x24/" target="_blank" rel="noopener">Seagate Exos 24 HDs</a> are 3.5&quot; 24T HDDs. This would allow us to triple the storage on existing servers. Each HDD is about $500, so upgrading one 2U machine would be about $6,000 and have storage of 220T or so.</p><ul> <li><em>Pros</em>: Reuses existing hardware we already have. Upgrades can be done a machine at a time. Fairly low price</li> <li><em>Cons</em>: Will existing units handle 24T drives? What’s the rebuild time on drive failure look like? It’s almost a day for 8T drives already, so possibly nearly a week for a failed 24T drive? Is there enough IO performance to handle daily backups at capacity?</li> </ul><h3 id="3-upgrade-to-new-hardware" tabindex="-1">3. Upgrade to new hardware</h3><p>As we know, SSDs are denser (2.5&quot; -&gt; 24 per 2U vs 3.5&quot; -&gt; 12 per 2U), more reliable, and now higher capacity - <a href="https://www.solidigm.com/products/data-center/d5/p5336.html#form=U.2%2015mm&amp;cap=61.44TB" target="_blank" rel="noopener">up to 61T per 2.5&quot; drive</a>. A single 2U server with 24 x 61T drives with 2 x 12 RAIDz2 = 1220T. Each drive is <a href="https://www.newegg.com/solidigm-61-44tb-d5-p5336/p/N82E16820318031" target="_blank" rel="noopener">about $7k</a> right now, prices fluctuate. So all up 24 x $7k = $168k + ~$20k server =~ $190k for &gt; 1000T storage one-time cost.</p><ul> <li><em>Pros</em>: <strong>Much</strong> higher sequential and random IO than HDDs will ever have. Price &lt; 1 year of standard S3 storage. Internal to our WAN, no bandwidth costs and very low latency. No new development required, existing backup system will just work. Consolidate on single 2U platform for all storage (cyrus, db, backups) and SSD for all storage. Significant space and power savings over existing HDD based servers</li> <li><em>Cons</em>: Greater up front cost. Still need to predict and buy more servers as backups grow.</li> </ul><p>One thing you don’t see in this calculation is datacenter space, power, cooling, etc. The reason is that compared to the amortised yearly cost of a storage server like this, these are actually reasonably minimal these days, on the order of $3000/2U/year. Calculating person time is harder. We have a lot of home built automation systems that mean installing and running one more server has minimal marginal cost.</p><h3 id="result" tabindex="-1">Result</h3><p>We ended up going with the the new 2U servers option:</p><p><picture><source type="image/webp" srcset="/assets/images/nvme-imap-servers-AqR6yL3DlW-375.webp 375w, /assets/images/nvme-imap-servers-AqR6yL3DlW-750.webp 750w, /assets/images/nvme-imap-servers-AqR6yL3DlW-1500.webp 1500w" sizes="(max-width: 425px) 375px, 750px"><img alt="NVME IMAP Servers" loading="lazy" decoding="async" src="/assets/images/nvme-imap-servers-AqR6yL3DlW-375.png" width="1500" height="559" srcset="/assets/images/nvme-imap-servers-AqR6yL3DlW-375.png 375w, /assets/images/nvme-imap-servers-AqR6yL3DlW-750.png 750w, /assets/images/nvme-imap-servers-AqR6yL3DlW-1500.png 1500w" sizes="(max-width: 425px) 375px, 750px"></picture></p><ul> <li>The 2U AMD NVMe platform with ZFS is a platform we have experience with already</li> <li>SSDs are much more reliable and much higher IO compared to HDDs</li> <li>No uncertainty around super large HDDs, RAID controllers, rebuild times, shuffling data around, etc.</li> <li>Significant space and power saving over existing HDD based servers</li> <li>No new development required, can use existing backup system and code</li> <li>Long expected hardware lifetime, controlled upfront cost, can depreciate hardware cost</li> </ul><p>So far this has worked out very well. The machines have bonded 25Gbps networks and when filling them from scratch we were able to saturate the network links streaming around 5Gbytes/second of data from our IMAP servers, compressing and writing it all down to a RAIDz2 zstd-3 compressed ZFS dataset.</p><h2 id="conclusion" tabindex="-1">Conclusion</h2><p>Running your own hardware might not be for everyone and has distinct tradeoffs. But when you have the experience and the knowledge of how you expect to scale, the cost improvements can be significant.</p></content> 57 </entry><entry> 58 <title>Dec 21: Fastmail in a box</title> 59 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/fastmail-in-a-box/' /> 60 <id>https://www.fastmail.com/blog/fastmail-in-a-box/</id> 61 <updated>2024-12-21T00:00:01Z</updated><author> 62 <name>Andrew Davis</name> 63 </author><content xml:lang='en' type='html'><p>This is the twenty-first post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/how-fastmail-uses-fastmail/">Dec 20: How Fastmail uses Fastmail!</a>. The next post is <a href="/blog/why-we-use-our-own-hardware/">Dec 22: Why we use our own hardware at Fastmail</a>.</p><p>They say everybody has a testing environment. Some people are just lucky enough enough to have a separate environment for production. At Fastmail, every staff member can get their own isolated testing and development sandbox. We call this Fastmail-In-A-Box, or more commonly just “fminabox”.</p><p>Like many technologists, I learn most by fiddling with things, often breaking them along the way and putting them back together again. With fminabox, we give everyone their own world to break apart and put back together, risk free. This makes it an invaluable place for new hires to cut their teeth, and for existing staff to come up to speed in an area of Fastmail’s stack they haven’t worked on before.</p><p>Fminabox is a complete Fastmail deployment on a single host. This includes Cyrus for IMAP storage, Postfix for incoming and outgoing mail, MySQL for non-mail data, our JMAP web API, and all the frontend assets. It also runs the ancillary services we use to monitor Fastmail such as Prometheus, all managed by the same configuration and service management system we use in production. This allows for fast, iterative development with very little waiting time between making a change and seeing the effect, while eliminating most “it worked on my machine” bugs.</p><p>We use <a href="https://www.packer.io/" target="_blank" rel="noopener">Hashicorp Packer</a> to create fminabox, following the provisioning scripts we use in production as closely as possible. Who hasn’t made a change to a system where it works going from state N → N+1, but then discovered weeks later that it’s broken when bootstrapping from nothing? Each night we build a new image from scratch. This allows us to catch those types of failures, and to do so while the changes are still front-of-mind in the developers that made them.</p><p>Any staff member can tell our chatbot <a href="https://github.com/fastmail/Synergy" target="_blank" rel="noopener">Synergy</a> to <code>box create</code>, and Synergy will handle provisioning a VM in the cloud, set up DNS, and provide VPN configuration upon request. Fastmail continues to eschew the public cloud in favour of our own hardware to run our product, but it turns out the public cloud is really useful for creating test environments.</p><p>Fminabox is also a key part of our testing workflow. Fastmail has thousands of tests, from simple sanity compile checks to complex integration tests between systems. We use fminabox with our CI/CD pipeline so every change is automatically tested before it is merged. This was the ultimate progression from developers just running a handful of tests manually, to overnight runs, to fully integrated continuous testing.</p><p>As new needs arise, we continue to evolve the infrastructure. A few years ago I was making an improvement to our tooling that balances users between machines in our <a href="https://www.fastmail.com/blog/building-a-backup-system-for-cyrus/" target="_blank" rel="noopener">Cyrus backup system</a>. At the time, fminabox only had a single target that all users were backed up to, so my first step was to add support for multiple backup targets. Only then did I feel comfortable that I could properly test any changes to the tooling.</p><p>I’m not the only user, so I asked some other Fastmail staff members “what’s your favourite feature of fminabox?”, and here’s what they had to say.</p><blockquote> <p>Fastmail is able to send via externally authenticated submission via OAuth, but Fastmail is also an OAuth provider and provides authenticated SMTP submission via OAuth. We were able to update our test suite to do full end-to-end OAuth authentication with ourself, send an email back to ourselves, and see that this entire path works.</p> </blockquote><p>—Rob Mueller, CTO</p><blockquote> <p>The best thing about fminabox is that it’s cheap and disposable. If I mess it up, I throw it away and make a new one and act like nothing happened. (The previous solution took hours to create a new box.)</p> </blockquote><p>—Ricardo Signes, Head of Special Projects</p><blockquote> <p>Although it only takes 5 minutes to setup, inaboxes provide a fully contained sandbox that includes all of our code ready to test. Minutes to build, seconds to tear down.</p> </blockquote><p>—Marcus Love, System Engineer</p><p>And that’s the story of fminabox. It isn’t perfect but it’s pretty damn good and it helps enable my colleagues to get their work done.</p></content> 64 </entry><entry> 65 <title>Dec 20: How Fastmail uses Fastmail!</title> 66 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/how-fastmail-uses-fastmail/' /> 67 <id>https://www.fastmail.com/blog/how-fastmail-uses-fastmail/</id> 68 <updated>2024-12-20T00:00:01Z</updated><author> 69 <name>Anju Manohar</name> 70 </author><content xml:lang='en' type='html'><p>This is the twentieth post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/offline-mail-storage/">Dec 19: Building offline: mail storage</a>. The next post is <a href="/blog/fastmail-in-a-box/">Dec 21: Fastmail in a box</a>.</p><p>At Fastmail, our features are designed to make email management seamless and efficient. But how do our own staff use these tools in their daily lives? While all the features are incredibly helpful, there will always be some personal favorites for one.</p><p>We asked team members to share their favorite Fastmail features and how they’ve customized them to fit their workflows. As support staff, we’re familiar with every feature, but each of us uses them in unique ways—and you might just discover a hidden gem in this article by seeing how we put them to work. And here’s what they had to say.</p><hr><p><strong>Vysakh: Mastering Inbox Zero and Organization</strong></p><p><em>“I enjoy maintaining an Inbox Zero approach, so I organize my emails into folders for various services.”</em></p><p>Vysakh takes organization to the next level with <a href="https://www.fastmail.help/hc/en-us/articles/1500000280301-Setting-up-and-using-folders" target="_blank" rel="noopener"><strong>Folders</strong></a> and <a href="https://www.fastmail.help/hc/en-us/articles/1500000278122-Mail-rules" target="_blank" rel="noopener"><strong>Rules</strong></a>. He creates folders for specific categories like “Bank” (with subfolders for each bank) and “Purchases” (with subfolders for each website he purchases from like Amazon and Flipkart).</p><p>He also raves about <a href="https://www.fastmail.help/hc/en-us/articles/360060591213-Searching-your-mail#:~:text=Saved%20searches,for%20the%20same%20thing%20later." target="_blank" rel="noopener"><strong>Saved Searches</strong></a>, calling them an <em>“underrated feature”</em>:</p><ul> <li>A saved search for <strong>Unread Emails</strong> allows him to see all unread messages across folders.</li> <li>Another for <strong>Emails Delivered Today</strong> is perfect for quickly finding new emails, even if they’re in Spam or Trash.</li> </ul><p>For self-organization, Vysakh uses <a href="https://www.fastmail.help/hc/en-us/articles/360060591053-Plus-addressing-and-subdomain-addressing" target="_blank" rel="noopener"><strong>Plus Addressing</strong></a> to save important documents by emailing them to himself with a folder-specific alias like <code>username+Docs@fastmail.tld</code>. He adds, <em>“Fastmail supports searching inside attachments, so finding them later is super easy.”</em></p><p>Vysakh also highlights the efficiency of <strong>Keyboard Shortcuts</strong> for email navigation and <strong>Customizable Notification Actions</strong> in the Android app, which let him manage emails without opening the app.</p><p>I must say, he’s truly a pro-user of all our power features!</p><hr><p><strong>Merlin: Timing and Efficiency Made Easy</strong></p><p>Merlin’s favorite is <a href="https://www.fastmail.help/hc/en-us/articles/4686576659727-Scheduled-Send" target="_blank" rel="noopener"><strong>Schedule Send</strong></a>, which allows drafting emails at their convenience and sending them at the perfect time. I couldn’t agree more—it’s like having a personal assistant keeping your inbox in check!</p><p>She goes on to say, <em>“My second favorite is Snooze—it lets me come back to emails when I have time to act on them.”</em> Smart, right? Fastmail features truly help you work smarter, not harder.</p><p>Merlin also loves <strong>Mail rules</strong> to keep her emails organized effortlessly and appreciates their simplicity, saying, <em>&quot;It is a cinch to use even for beginners.”</em></p><hr><p><strong>Maya: A Domain for Every Interest</strong></p><p>Maya’s love for email personalization shines through her extensive use of <a href="https://www.fastmail.help/hc/en-us/articles/360058753394-Custom-domains-with-Fastmail" target="_blank" rel="noopener"><strong>Domains</strong></a> and <a href="https://www.fastmail.help/hc/en-us/articles/360060591073-How-to-set-up-aliases" target="_blank" rel="noopener"><strong>Aliases</strong></a>:</p><p><em>“I have a whole bunch of domains just for fun—some professional, some for hobbies like photo essays and writing.”</em></p><p>Her go-to feature is <a href="https://www.fastmail.help/hc/en-us/articles/4406536368911-Masked-Email" target="_blank" rel="noopener"><strong>Masked Email</strong></a> combined with a custom domain. Whenever she signs up for a service, she creates a unique masked address, assigning each one to a folder. This setup lets her instantly identify breaches, unsubscribe from unwanted notifications, and organize her inbox for easy prioritization.</p><p>Maya also uses:</p><ul> <li><strong>Pins</strong> to flag important emails in folders.</li> <li><strong>Snooze</strong> for emails she wants to revisit later without cluttering her inbox.</li> <li><strong>Tamper-Proof Retention</strong>, ensuring emails she’s deleted can still be recovered if needed.</li> </ul><p>She admits with a laugh, <em>“It’s probably not something a non-business user like me should need, but it’s handy!”</em></p><hr><p><strong>Thu: Seamless Sharing and Catchall Convenience</strong></p><p>Thu finds <a href="https://www.fastmail.help/hc/en-us/articles/1500000277942-Catch-all-wildcard-aliases" target="_blank" rel="noopener"><strong>Catchall aliases</strong></a> and <a href="https://www.fastmail.help/hc/en-us/articles/360060590733-Sharing-mail" target="_blank" rel="noopener"><strong>Mail sharing</strong></a> to be game-changers. <em>“The catchall is super handy for creating new email addresses anytime I want”.</em> She’s even set up a Fastmail account for her partner with shared aliases, mail, and calendars. Excitedly, she adds, <em>“The sharing function works very well with integration to Apple devices and Gmail account,”</em> making collaboration a breeze. You can almost feel her joy through those words—don’t you? I bet you do!</p><hr><p><strong>Leslie: Staying Organized with Pins</strong></p><p>Leslie swears by the <a href="https://www.fastmail.help/hc/en-us/articles/1500000280341-Pin-important-messages" target="_blank" rel="noopener"><strong>Pin</strong></a> and <strong>Keep pinned on top</strong> features to stay on top of essential emails. <em>“I use this on my staff account quite a lot because we get various emails with a bit of information in them. When I go through my emails every morning, I will pin the emails that I want to go back and read and also emails about support trends that are occurring. This allows me to quickly refer back to the emails since I’m seeing them at the top of the list in my inbox.”</em> Leslie clearly has a knack for using this feature to keep important information right at her fingertips—guess I know who to turn to for a quick reference!</p><hr><p><strong>Yassar: Power User of Labels and Searches</strong></p><p>For Yassar, <a href="https://www.fastmail.help/hc/en-us/articles/360058753554-Setting-up-and-using-labels" target="_blank" rel="noopener"><strong>Labels</strong></a> are indispensable. <em>“I love organizing everything in my mailbox, and I can’t imagine working without labels!”</em> If you’re a fan of staying organized, you’re probably nodding in agreement right now.</p><p>But Yassar doesn’t stop there, his other go-to features include:</p><ul> <li><strong>Masked Emails</strong> for secure online shopping.</li> <li><strong>Aliases</strong> to classify emails and route them to specific folders such as <code>documents@mydomain.com</code> or <code>gmail@mydomain.com</code>.</li> <li><strong>Saved Search</strong> for quick filtering of emails without creating a label for everything.</li> </ul><p>Yassar’s approach combines security, organization, and speed, showing how powerful these tools can be for anyone looking to optimize their email experience!</p><hr><p><strong>Jed: Streamlining Inbox Management with Keyboard Shortcuts</strong></p><p>Jed relies heavily on <a href="https://www.fastmail.help/hc/en-us/articles/360058753534-Keyboard-shortcuts" target="_blank" rel="noopener"><strong>Keyboard Shortcuts</strong></a> as one of his favorite Fastmail features:</p><p><em>“As someone who lets a lot of email build up before I get around to sorting them, I like how quickly I can label, archive, and delete large numbers of messages using these shortcuts.”</em></p><p>Wow! That’s brilliant—such a simple feature, yet so powerful! He’s turned what could be a daunting task into a seamless process. I’m definitely going to start using these more, and you should too!</p><hr><p><strong>Jess: Memos, Snoozing, and Smart Searching</strong></p><p>Jess praises <a href="https://www.fastmail.help/hc/en-us/articles/1500001969861-Conversations#:~:text=conversations%20help%20page.-,Memos,-Use%20memos%20to" target="_blank" rel="noopener"><strong>Memos</strong></a> for saving key details from lengthy emails. With a grin, she shares, <em>“I just used it to note a coupon hidden in a long marketing email.”</em> Who else can relate? I know I can—if you’re like me, always hunting for those elusive discount codes, Memos are here to save the day!</p><p>She’s also a big fan of the <a href="https://www.fastmail.help/hc/en-us/articles/360058753634-Snoozing-mail" target="_blank" rel="noopener"><strong>Snooze</strong></a> feature, which she frequently uses for bills or shipping notifications that need attention later.</p><p>Lastly, Jess makes extensive use of <strong>Search</strong> to streamline her inbox. The <code>is:unread</code> search helps her capture all unread emails across labels, enabling her to quickly sort through them and get closer to Inbox Zero. That’s an amazing tip—you’ll get a unified inbox with all unread emails in one place!</p><hr><p><strong>Aric: Sending at the Perfect Time</strong></p><p><strong>Scheduled send</strong> is one of the features Aric relies on the most, and it’s no surprise why, as he shares: “<em>Working with teams across the globe and working on a non-traditional schedule, I often find myself sending mail outside of people’s general working hours. Using scheduled send allows me to send mail so that my message is one of the first things the recipient sees when they check their inbox.&quot;</em> Impressive—that explains why his emails always seem to land at just the right time!</p><hr><p>It’s been such a joy hearing from our team members about their favorite features. Inspired by them, how can I not share mine? I’m certainly not stepping back—here are my favorites.</p><p>One feature I find myself using a lot is the <a href="https://www.fastmail.help/hc/en-us/articles/360058752514-Logged-in-sessions#:~:text=wish%20to%20end.-,View%20all%20logins%20in%20the%20last%204%20weeks,-At%20the%20bottom" target="_blank" rel="noopener"><strong>Login Log</strong></a>, which is crucial for maintaining the privacy and security of my account. It lets me instantly review login activities whenever I spot something suspicious. The interface is simple and intuitive, showing logins from the Fastmail web UI, Fastmail app, and third-party apps, along with any failed login attempts. This feature gives me the confidence to monitor and act swiftly when needed—no more panic attacks when something looks off!</p><p>Next up is the <a href="https://www.fastmail.help/hc/en-us/articles/360060591213-Searching-your-mail#:~:text=searches%20with%20syntax.-,Advanced%20search,-If%20you%27re%20having" target="_blank" rel="noopener"><strong>Advanced Search</strong></a>, a powerful yet user-friendly tool. For someone new to search tools and struggling to remember search syntax, this feature would be a real lifesaver. I use it all the time to refine my searches by selecting criteria from the available fields—whether it’s finding emails with specific attachment types, emails within a date range, or so much more. It’s fast, easy, and incredibly efficient! If you haven’t tried it yet, give it a go today—you’ll be amazed at how simple it is!</p><p>Then comes the <a href="https://www.fastmail.help/hc/en-us/articles/10351615144335-Passkeys" target="_blank" rel="noopener"><strong>Passkey</strong></a> feature making my life easier and logging into my account across different devices both simpler and more secure.</p><p>Honestly, I love all the features! They’ve taken me—once a self-proclaimed email management avoider (guilty as charged!)—and turned me into a full-fledged Inbox enthusiast.</p><hr><p>A big thank you to my team for sharing their insights. It’s been great learning how everyone makes the most of these tools!</p><p>That wraps up an exciting dive into these amazing features and creative ways to use them—I hope you’ve found some inspiration to explore and make them your own!</p><p>Still not on Fastmail? Now’s the perfect time! Get your whole family on board with the Family plan and enjoy an <a href="https://app.fastmail.com/signup/?discount=MjAsMTIsMTczNTY0NzQ0MCxhbWFub2hhciwsZjMyN2RlNGU4ZDVmZTM4MmVmNDM3YzI2YzRhNDFhY2EyYmZmMDA4MA" target="_blank" rel="noopener">exclusive 20% off your first year</a>—don’t miss out!</p></content> 71 </entry><entry> 72 <title>Dec 19: Building offline: mail storage</title> 73 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/offline-mail-storage/' /> 74 <id>https://www.fastmail.com/blog/offline-mail-storage/</id> 75 <updated>2024-12-19T00:00:01Z</updated><author> 76 <name>Neil Jenkins</name> 77 </author><content xml:lang='en' type='html'><p>This is the nineteenth post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/offline-sync/">Dec 18: Building offline: syncing changes back to the server</a>. The next post is <a href="/blog/how-fastmail-uses-fastmail/">Dec 20: How Fastmail uses Fastmail!</a>.</p><p>Yesterday, we looked at <a href="/blog/offline-sync/">how we store changes you make offline</a> so we can accurately and efficiently sync them back to the server when you come online. Today, we’ll discuss why email is special, and what else we do to make this super fast, with support for full-text search offline.</p><h2 id="why-offline-email-is-hard" tabindex="-1">Why offline email is hard</h2><p>As discussed earlier, because we use JMAP for all of our APIs, once we can implement generic offline support and have it work for everything (currently 56 data types and counting in our app!). However, mail is special. And the reason it’s special is purely the volume of data.</p><p>Most web apps severely underestimate how small their data is. In almost all cases, you will be more efficient and way faster to just suck it all into memory and do a linear filter pass whenever you need to query it. This is the difference between response as-you-type autocomplete and frustrating loading spinners on each key stroke. Even for users with 10,000 contacts this is only a few megabytes of data — perfectly cacheable.</p><p>Email is different though. We have users with millions of messages. Even with attachments handled separately in JMAP, each message could have hundreds of kilobytes of HTML as the body. But we expect opening a mailbox to load a listing pretty much instantly, and searches to be fast too. To make this work, we have to add a number of tricks to our standard offline approach.</p><h2 id="splitting-the-data" tabindex="-1">Splitting the data</h2><p>The first trick is to split the data into two separate <a href="https://developer.mozilla.org/en-US/docs/Web/API/IDBObjectStore" target="_blank" rel="noopener">object stores</a>:</p><ol> <li> <p><strong>EmailMetadata</strong>: this stores just the data that’s not parsed from the email content, like the id, thread id, keywords it has, and mailboxes it’s in. This keeps it small, but crucially also contains all the mutable data. This is treated like our standard JMAP object store for a data type.</p> </li> <li> <p><strong>EmailContent</strong>: this stores the email content; who it was sent from/to, the subject, body, list of attachments (but not the attachment data itself) etc.</p> </li> </ol><p>Due to the volume of data, we can’t load everything at once. We page in the data in stages instead:</p><ol> <li>We fetch a list of just the ids and create placeholder entries in the EmailMetadata object store.</li> <li>We page in the metadata and basic headers (like to/from/subject) for all messages in batches. This gives us everything we need to show the listing for any folder or label.</li> <li>We page in the body for pinned and recent messages, or everything if the user has selected this option in settings, again in batches.</li> </ol><p>This split is useful, because for most queries we can get away with just loading the metadata into memory, not the content. This is a big saving in time and memory when deserialising the objects from the underlying datastore.</p><h2 id="efficient-mailbox-querying" tabindex="-1">Efficient mailbox querying</h2><p>A linear pass through all the metadata is surprisingly tractable, even for large mailboxes, however it’s slower than we want for common queries (like opening your inbox). This is where we introduce a couple of extra custom indexes — separate object stores we are careful to update in lock step with any changes to our data.</p><p>The first of these is <strong>EmailMailboxes</strong>. This stores an entry for each addition or removal of a message from a folder/label, allowing us to both very efficiently compute the list of messages/conversations in a particular mailbox, and also calculate a delta update to the query when making changes.</p><p>The key for this object store is:</p><pre class="language-javascript"><code class="language-javascript"><span class="punctuation token">[</span><span class="constant token">MAILBOX_ID</span><span class="punctuation token">,</span> <span class="constant token">REMOVED_MODSEQ</span><span class="punctuation token">,</span> <span class="constant token">ADDED_MODSEQ</span><span class="punctuation token">]</span><span class="punctuation token">;</span></code></pre><p>The values look like:</p><pre class="language-javascript"><code class="language-javascript"><span class="punctuation token">[</span><span class="constant token">EMAIL_ID</span><span class="punctuation token">,</span> <span class="constant token">THREAD_ID</span><span class="punctuation token">,</span> <span class="constant token">DATE</span><span class="punctuation token">,</span> <span class="constant token">IS_UNREAD</span><span class="punctuation token">]</span><span class="punctuation token">;</span></code></pre><p>Whenever a message is added to a mailbox, a new entry is created. <code>ADDED_MODSEQ</code> is the current “updated” moseq of the message, and <code>REMOVED_MODSEQ</code> is 0.</p><p>If the message is removed from the mailbox, the old entry is deleted, and a new one added with the same <code>ADDED_MODSEQ</code>, but <code>REMOVED_MODSEQ</code> set to the new “updated” modseq of the message.</p><p>From this, we can quickly get the list of current messages in a particular mailbox by doing a range query for entries with keys that start: <code>[MAILBOX_ID, 0]</code>. The values include the date and thread id, allowing us to do the most common sort, and remove duplicates for the same thread id, without having to even fetch the metadata objects for the emails.</p><h2 id="delta-query-updates" tabindex="-1">Delta query updates</h2><p>JMAP has a way for a client to <a href="https://www.rfc-editor.org/rfc/rfc8620.html#section-5.6" target="_blank" rel="noopener">ask for what’s changed in a query</a>. This allows it to more efficiently update its local store and uses less bandwidth. With the EmailMailboxes index, we can also implement this. First we fetch the entries for the current messages as before, but then we also fetch the entries for messages that have been removed since our last state (this is a range query between <code>[MAILBOX_ID, sinceModSeq + 1]</code> and <code>[MAILBOX_ID, max_int]</code>). We sort these entries together according to the sort order the user has requested, normally date descending:</p><pre class="language-javascript"><code class="language-javascript">mailboxRecords<span class="punctuation token">.</span><span class="function token">sort</span><span class="punctuation token">(</span> 78 <span class="punctuation token">(</span><span class="parameter token">a<span class="punctuation token">,</span> b</span><span class="punctuation token">)</span> <span class="operator token">=></span> 79 b<span class="punctuation token">[</span><span class="constant token">DATE</span><span class="punctuation token">]</span> <span class="operator token">-</span> a<span class="punctuation token">[</span><span class="constant token">DATE</span><span class="punctuation token">]</span> <span class="operator token">||</span> 80 <span class="punctuation token">(</span>a<span class="punctuation token">[</span><span class="constant token">EMAIL_ID</span><span class="punctuation token">]</span> <span class="operator token">&lt;</span> b<span class="punctuation token">[</span><span class="constant token">EMAIL_ID</span><span class="punctuation token">]</span> <span class="operator token">?</span> <span class="number token">1</span> <span class="operator token">:</span> a<span class="punctuation token">[</span><span class="constant token">EMAIL_ID</span><span class="punctuation token">]</span> <span class="operator token">></span> b<span class="punctuation token">[</span><span class="constant token">EMAIL_ID</span><span class="punctuation token">]</span> <span class="operator token">?</span> <span class="operator token">-</span><span class="number token">1</span> <span class="operator token">:</span> <span class="number token">0</span><span class="punctuation token">)</span> <span class="operator token">||</span> 81 a<span class="punctuation token">[</span><span class="constant token">ADDED_MODSEQ</span><span class="punctuation token">]</span> <span class="operator token">-</span> b<span class="punctuation token">[</span><span class="constant token">ADDED_MODSEQ</span><span class="punctuation token">]</span><span class="punctuation token">,</span> 82 <span class="punctuation token">)</span><span class="punctuation token">;</span></code></pre><p>Then we can iterate through to calculate what has been added or removed from the query, like so. (“Exemplar” is our term for the email that’s representing a thread when <a href="https://www.rfc-editor.org/rfc/rfc8621.html#section-4.4.3" target="_blank" rel="noopener">the “collapseThreads” argument</a> is true.)</p><pre class="language-javascript"><code class="language-javascript"><span class="keyword token">let</span> index <span class="operator token">=</span> <span class="operator token">-</span><span class="number token">1</span><span class="punctuation token">;</span> 83 <span class="keyword token">const</span> seenExemplar <span class="operator token">=</span> collapseThreads <span class="operator token">?</span> <span class="keyword token">new</span> <span class="class-name token">Set</span><span class="punctuation token">(</span><span class="punctuation token">)</span> <span class="operator token">:</span> <span class="keyword token">null</span><span class="punctuation token">;</span> 84 <span class="keyword token">const</span> seenOldExemplar <span class="operator token">=</span> collapseThreads <span class="operator token">?</span> <span class="keyword token">new</span> <span class="class-name token">Set</span><span class="punctuation token">(</span><span class="punctuation token">)</span> <span class="operator token">:</span> <span class="keyword token">null</span><span class="punctuation token">;</span> 85 <span class="keyword token">let</span> uptoHasBeenFound <span class="operator token">=</span> <span class="boolean token">false</span><span class="punctuation token">;</span> 86 <span class="keyword token">let</span> total <span class="operator token">=</span> <span class="number token">0</span><span class="punctuation token">;</span> 87 <span class="keyword token">const</span> added <span class="operator token">=</span> <span class="punctuation token">[</span><span class="punctuation token">]</span><span class="punctuation token">;</span> 88 <span class="keyword token">const</span> removed <span class="operator token">=</span> <span class="punctuation token">[</span><span class="punctuation token">]</span><span class="punctuation token">;</span> 89 <span class="keyword token">for</span> <span class="punctuation token">(</span><span class="keyword token">const</span> record <span class="keyword token">of</span> mailboxRecords<span class="punctuation token">)</span> <span class="punctuation token">{</span> 90 <span class="keyword token">const</span> isDeleted <span class="operator token">=</span> <span class="operator token">!</span><span class="operator token">!</span>record<span class="punctuation token">[</span><span class="constant token">REMOVED_MODSEQ</span><span class="punctuation token">]</span><span class="punctuation token">;</span> 91 <span class="comment token">// Created and deleted after our previous state? Ignore.</span> 92 <span class="keyword token">const</span> isNew <span class="operator token">=</span> record<span class="punctuation token">[</span><span class="constant token">ADDED_MODSEQ</span><span class="punctuation token">]</span> <span class="operator token">></span> sinceModSeq<span class="punctuation token">;</span> 93 <span class="keyword token">if</span> <span class="punctuation token">(</span>isNew <span class="operator token">&amp;&amp;</span> isDeleted<span class="punctuation token">)</span> <span class="punctuation token">{</span> 94 <span class="keyword token">continue</span><span class="punctuation token">;</span> 95 <span class="punctuation token">}</span> 96 97 <span class="comment token">// Is this message the current exemplar?</span> 98 <span class="keyword token">let</span> isNewExemplar <span class="operator token">=</span> <span class="boolean token">false</span><span class="punctuation token">;</span> 99 <span class="keyword token">let</span> isOldExemplar <span class="operator token">=</span> <span class="boolean token">false</span><span class="punctuation token">;</span> 100 <span class="keyword token">const</span> emailId <span class="operator token">=</span> record<span class="punctuation token">[</span><span class="constant token">EMAIL_ID</span><span class="punctuation token">]</span><span class="punctuation token">;</span> 101 <span class="keyword token">const</span> threadId <span class="operator token">=</span> record<span class="punctuation token">[</span><span class="constant token">THREAD_ID</span><span class="punctuation token">]</span><span class="punctuation token">;</span> 102 <span class="keyword token">if</span> <span class="punctuation token">(</span><span class="operator token">!</span>isDeleted <span class="operator token">&amp;&amp;</span> <span class="punctuation token">(</span><span class="operator token">!</span>collapseThreads <span class="operator token">||</span> <span class="operator token">!</span>seenExemplar<span class="punctuation token">.</span><span class="function token">has</span><span class="punctuation token">(</span>threadId<span class="punctuation token">)</span><span class="punctuation token">)</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 103 isNewExemplar <span class="operator token">=</span> <span class="boolean token">true</span><span class="punctuation token">;</span> 104 index <span class="operator token">+=</span> <span class="number token">1</span><span class="punctuation token">;</span> 105 total <span class="operator token">+=</span> <span class="number token">1</span><span class="punctuation token">;</span> 106 <span class="keyword token">if</span> <span class="punctuation token">(</span>collapseThreads<span class="punctuation token">)</span> <span class="punctuation token">{</span> 107 seenExemplar<span class="punctuation token">.</span><span class="function token">add</span><span class="punctuation token">(</span>threadId<span class="punctuation token">)</span><span class="punctuation token">;</span> 108 <span class="punctuation token">}</span> 109 <span class="punctuation token">}</span> 110 <span class="comment token">// Was this message an old exemplar?</span> 111 <span class="comment token">// 1. Must not have been added to mailbox after the client's state</span> 112 <span class="comment token">// 2. Must have been removed from mailbox before the client's state</span> 113 <span class="comment token">// 3. Must not have already found the old exemplar.</span> 114 <span class="keyword token">if</span> <span class="punctuation token">(</span><span class="operator token">!</span>isNew <span class="operator token">&amp;&amp;</span> <span class="punctuation token">(</span><span class="operator token">!</span>collapseThreads <span class="operator token">||</span> <span class="operator token">!</span>seenOldExemplar<span class="punctuation token">.</span><span class="function token">has</span><span class="punctuation token">(</span>threadId<span class="punctuation token">)</span><span class="punctuation token">)</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 115 isOldExemplar <span class="operator token">=</span> <span class="boolean token">true</span><span class="punctuation token">;</span> 116 <span class="keyword token">if</span> <span class="punctuation token">(</span>collapseThreads<span class="punctuation token">)</span> <span class="punctuation token">{</span> 117 seenOldExemplar<span class="punctuation token">.</span><span class="function token">add</span><span class="punctuation token">(</span>threadId<span class="punctuation token">)</span><span class="punctuation token">;</span> 118 <span class="punctuation token">}</span> 119 <span class="punctuation token">}</span> 120 121 <span class="keyword token">if</span> <span class="punctuation token">(</span>isOldExemplar <span class="operator token">&amp;&amp;</span> <span class="operator token">!</span>isNewExemplar<span class="punctuation token">)</span> <span class="punctuation token">{</span> 122 removed<span class="punctuation token">.</span><span class="function token">push</span><span class="punctuation token">(</span>emailId<span class="punctuation token">)</span><span class="punctuation token">;</span> 123 <span class="punctuation token">}</span> <span class="keyword token">else</span> <span class="keyword token">if</span> <span class="punctuation token">(</span><span class="operator token">!</span>isOldExemplar <span class="operator token">&amp;&amp;</span> isNewExemplar<span class="punctuation token">)</span> <span class="punctuation token">{</span> 124 <span class="comment token">// If the message has been moved out and back in again</span> 125 <span class="comment token">// we'll have separate mailbox records for added/removed</span> 126 <span class="comment token">// so not detect it's both the old and new exemplar;</span> 127 <span class="comment token">// check for that here.</span> 128 <span class="keyword token">const</span> removedIndex <span class="operator token">=</span> isMutableSort <span class="operator token">?</span> <span class="operator token">-</span><span class="number token">1</span> <span class="operator token">:</span> removed<span class="punctuation token">.</span><span class="function token">indexOf</span><span class="punctuation token">(</span>emailId<span class="punctuation token">)</span><span class="punctuation token">;</span> 129 <span class="keyword token">if</span> <span class="punctuation token">(</span>removedIndex <span class="operator token">></span> <span class="operator token">-</span><span class="number token">1</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 130 removed<span class="punctuation token">.</span><span class="function token">splice</span><span class="punctuation token">(</span>removedIndex<span class="punctuation token">,</span> <span class="number token">1</span><span class="punctuation token">)</span><span class="punctuation token">;</span> 131 <span class="punctuation token">}</span> <span class="keyword token">else</span> <span class="punctuation token">{</span> 132 added<span class="punctuation token">.</span><span class="function token">push</span><span class="punctuation token">(</span><span class="punctuation token">{</span> 133 index<span class="punctuation token">,</span> 134 <span class="literal-property property token">id</span><span class="operator token">:</span> emailId<span class="punctuation token">,</span> 135 <span class="punctuation token">}</span><span class="punctuation token">)</span><span class="punctuation token">;</span> 136 <span class="punctuation token">}</span> 137 <span class="punctuation token">}</span> 138 139 <span class="comment token">// Special case for mutable sorts (based on isFlagged/isUnread)</span> 140 <span class="keyword token">if</span> <span class="punctuation token">(</span>isMutableSort <span class="operator token">&amp;&amp;</span> isOldExemplar <span class="operator token">&amp;&amp;</span> isNewExemplar<span class="punctuation token">)</span> <span class="punctuation token">{</span> 141 <span class="comment token">// Has the isUnread/isFlagged status of the message/thread</span> 142 <span class="comment token">// (as appropriate) possibly changed since the client's state?</span> 143 <span class="comment token">// If so, we need to remove the exemplar from the client view</span> 144 <span class="comment token">// and add it back in at the correct position.</span> 145 <span class="keyword token">const</span> mayHaveMoved <span class="operator token">=</span> collapseThreads 146 <span class="operator token">?</span> threadChanged<span class="punctuation token">.</span><span class="function token">has</span><span class="punctuation token">(</span>threadId<span class="punctuation token">)</span> 147 <span class="operator token">:</span> emailChanged<span class="punctuation token">.</span><span class="function token">has</span><span class="punctuation token">(</span>emailId<span class="punctuation token">)</span><span class="punctuation token">;</span> 148 <span class="keyword token">if</span> <span class="punctuation token">(</span>mayHaveMoved<span class="punctuation token">)</span> <span class="punctuation token">{</span> 149 removed<span class="punctuation token">.</span><span class="function token">push</span><span class="punctuation token">(</span>emailId<span class="punctuation token">)</span><span class="punctuation token">;</span> 150 added<span class="punctuation token">.</span><span class="function token">push</span><span class="punctuation token">(</span><span class="punctuation token">{</span> 151 index<span class="punctuation token">,</span> 152 <span class="literal-property property token">id</span><span class="operator token">:</span> emailId<span class="punctuation token">,</span> 153 <span class="punctuation token">}</span><span class="punctuation token">)</span><span class="punctuation token">;</span> 154 <span class="punctuation token">}</span> 155 <span class="punctuation token">}</span> 156 <span class="comment token">// If this is the last message the client cares about, we can stop</span> 157 <span class="comment token">// here and just return what we've calculated so far. We already</span> 158 <span class="comment token">// know the total count for this message list as we keep it pre</span> 159 <span class="comment token">// calculated and cached in the Mailbox object.</span> 160 <span class="comment token">// However, if the sort is mutable we can't break early, as</span> 161 <span class="comment token">// messages may have moved from the region we care about to lower</span> 162 <span class="comment token">// down the list.</span> 163 <span class="keyword token">if</span> <span class="punctuation token">(</span><span class="operator token">!</span>isMutableSort <span class="operator token">&amp;&amp;</span> <span class="operator token">!</span>isNew <span class="operator token">&amp;&amp;</span> emailId <span class="operator token">===</span> upToId<span class="punctuation token">)</span> <span class="punctuation token">{</span> 164 uptoHasBeenFound <span class="operator token">=</span> <span class="boolean token">true</span><span class="punctuation token">;</span> 165 <span class="keyword token">break</span><span class="punctuation token">;</span> 166 <span class="punctuation token">}</span> 167 <span class="punctuation token">}</span></code></pre><h2 id="mail-search" tabindex="-1">Mail search</h2><p>Fastmail supports an <a href="https://www.fastmail.help/hc/en-us/articles/360060591213-Searching-your-mail" target="_blank" rel="noopener">extremely powerful set of search operators</a>, allowing for <a href="https://www.fastmail.com/features/search/" target="_blank" rel="noopener">fast, precise searching</a>. We support almost all of it offline, with a few caveats discussed below.</p><p>To make full-text search work and be performant, we need to build another index. If you have hundreds of thousands of messages, it would be unusably slow to scan through all of them looking for a word, phrase or email address.</p><p>Our index is stored in another <a href="https://developer.mozilla.org/en-US/docs/Web/API/IndexedDB_API" target="_blank" rel="noopener">IndexedDB</a> object store called <strong>EmailSearch</strong>. The key for each entry is <code>[token, emailId]</code>. The token is usually a word or other sequence of letters and numbers extracted from the email. We also have special token variations to represent a list-id or email addresses found in the headers. We create an entry in EmailSearch for each such token we find in the email. The value encodes where the token was found (e.g. in the <code>To</code> header, or the message body), and the index(es) of the token so we can do <a href="https://en.wikipedia.org/wiki/Phrase_search" target="_blank" rel="noopener">phrase searches</a>.</p><p>We decided to index the content on the device, rather than download the indexes from the server. This ensured our search index would be completely in sync with the cached messages you have on your device, and we could index and make searchable messages and memos you wrote while you were offline.</p><p>However, this does mean the offline search works a little differently to our server-based search, so may return slightly different results (although we think both will do a great job in most cases). In particular:</p><ul> <li>Our offline search doesn’t index any text inside attachments. When online you can search for content in attached PDFs, spreadsheets, and other documents.</li> <li>Our offline search doesn’t do <a href="https://en.wikipedia.org/wiki/Stemming" target="_blank" rel="noopener">stemming</a>. Stemming tries to reduce a word to its common root, so if you search in English for <code>bus</code> you would also match emails containing <code>buses</code>, but not <code>business</code>. Stemming requires language analysis of the email content and custom stemming algorithms for each language, and we decided the extra complexity and code download size was not currently worth it for our offline search. Instead, our offline search does prefix matching by default, so <code>bus</code> will still match <code>buses</code> but also <code>business</code>. Of course, if you wrap the term in quotes (like <code>&quot;bus&quot;</code>) it will only look for exact matches, just like with server-based search.</li> </ul><p>And of course, the search index will only contain messages you have downloaded for offline, which might not be everything in your account. We therefore try to do a search on the server first and only fallback to the local search if you are offline.</p><h2 id="search-tokenisation" tabindex="-1">Search tokenisation</h2><p>To create our index we have to be able to extract the tokens from a sequence of text. We have users around the world, so we knew we had to handle multilingual text and scripts. In the end, we settled on a simple but effective tokenisation algorithm:</p><ol> <li>We normalise the string into Unicode <a href="https://en.wikipedia.org/wiki/Unicode_equivalence#Normal_forms" target="_blank" rel="noopener">NFKD normal form</a>. This will decompose diacritics to make it easy to strip them, and replace various variations of letters and numbers (such as typographic ligatures, or subscript numbers) with the baseline equivalent.</li> <li>We divide the string into segments according to the <a href="https://www.unicode.org/reports/tr29/#Word_Boundaries" target="_blank" rel="noopener">Unicode text segmentation word boundary algorithm</a>.</li> <li>For each segment, we apply the full <a href="https://www.unicode.org/Public/16.0.0/ucd/CaseFolding.txt" target="_blank" rel="noopener">Unicode case folding substitutions</a> (for example, this will replace uppercase letters with lowercase for Latin text), then we strip every code point that’s not categorised by Unicode as a number, letter, joining punctuation, or emoji.</li> </ol><p>If we have anything left, that’s our token. So to give an example, supposing we had the text:</p><pre><code>The café is über cheap — only $3.60 a ☕️!! 168 </code></pre><p>We would end up with the following tokens:</p><pre><code>the 169 cafe 170 is 171 uber 172 cheap 173 only 174 360 175 a 176 ☕️ 177 </code></pre><h2 id="wrapping-it-up" tabindex="-1">Wrapping it up</h2><p>We now have the indexes we need for fast, precise search. There’s still a lot of work involved in putting it all together though! When you search for something complex like <code>in:inbox from:@example.com (is:pinned OR &quot;very important&quot;)</code>, we analyse the query to work out which indexes to use and efficiently combine them to compute the results. The speed will depend on how much mail you have—and how fast your device is!—but we believe it lives up to the Fastmail promise of great search everywhere.</p><p>There’s so much interesting tech behind our offline support, but for now I need to stop writing. If you’ve read all of this mini series on how we are making our app work offline: thank you, and I hope you found it interesting! Please give the beta a go, and let us know any feedback you might have. We’re excited to finish polishing this highly requested feature and we hope to ship it to everyone early in the new year.</p></content> 178 </entry><entry> 179 <title>Dec 18: Building offline: syncing changes back to the server</title> 180 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/offline-sync/' /> 181 <id>https://www.fastmail.com/blog/offline-sync/</id> 182 <updated>2024-12-18T00:00:01Z</updated><author> 183 <name>Neil Jenkins</name> 184 </author><content xml:lang='en' type='html'><p>This is the eighteenth post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/offline-architecture/">Dec 17: Building offline: general architecture</a>. The next post is <a href="/blog/offline-mail-storage/">Dec 19: Building offline: mail storage</a>.</p><p>Yesterday, we looked at <a href="/blog/offline-architecture/">how our offline caching layer fits into our app</a>, and the way it stores data to efficiently respond to JMAP requests. Today, we’ll dive into how it keeps track of changes the user makes while offline, so it can reconcile this with the server.</p><h2 id="keeping-track-of-changes" tabindex="-1">Keeping track of changes</h2><p>When a client makes a change offline, we update our local cache and have to keep track of it so we can sync that change back to the server when we come online. There are two main approaches you could take:</p><ol> <li>You keep a time-ordered log of every change, then replay the log against the server. One record may appear multiple times in the log if it has multiple modifications applied.</li> <li>You keep a set of created/updated/destroyed records, along with the current server value. Each record can only appear once, in at most one of these categories. You calculate the difference between the server state and the current state to update the server.</li> </ol><p>The benefit of the first approach is it ensures we maintain any ordering dependencies. The benefit of the second approach is it’s more efficient in terms of both storage and synchronisation speed when there are multiple changes made to the same record.</p><p>The Fastmail offline cache uses a hybrid of these approaches to try to get the best of both worlds:</p><ul> <li>A log stores (in order) the <code>[data type, account id, id]</code> of any changes, along with what type of change this is (create/update/destroy).</li> <li>The record itself stores the last known server state if it’s been updated, stored efficiently as a patch to get back to the server state from the updated state.</li> <li>If the record is updated a second time, it: <ul> <li>stays in its current position in the log if not yet present on the server (this is a create); or</li> <li>moves to the end of the log (remove the old entry and add a new one) if it already exists on the server (this is an update/destroy); or</li> <li>is removed entirely from the log if the change reverted it back to the last-known server state.</li> </ul> </li> </ul><p>If we’re updating a record that’s not yet been created on the server, we may have to do an update as well as a create, due to an ordering problem. For example, suppose you do the following:</p><ol type="a"> <li>Create Mailbox X</li> <li>Create Emails A &amp; B in Mailbox X</li> <li>Create Mailbox Y</li> <li>Move Mailbox X to be a child of Y</li> <li>Move Email A to be in Mailbox Y</li> </ol><p>You can’t move (a) later because (b) depends on it. You can’t move (d) earlier because it depends on (c). So if we update a record that’s not yet been created on the server, and we set a property that includes a local id (i.e., it references another object that’s been created locally but not yet synced to the server), we add it as a patch and apply it as an update later.</p><p>When loading data from the server, we do not need to look for an entry in the log of changes still to sync. We can just update the server state in the record. If the change is now inert, we’ll delete it from the log when we go to sync it.</p><p>For example, suppose we have a mailbox, id <code>1</code>, with two messages in it, ids <code>A</code> &amp; <code>B</code>, and the user does the following (contrived) actions:</p><ul> <li>Creates a new mailbox: <code>2</code></li> <li>Creates a new child mailbox of that: <code>3</code></li> <li>Moves A and B into mailbox <code>3</code></li> <li>Marks B as read</li> <li>Moves A back to its original mailbox.</li> <li>Renames mailbox <code>2</code>.</li> </ul><p>Our log will end up looking like this:</p><pre><code> [Mailbox, &quot;#2&quot;, CREATE] 185 [Mailbox, &quot;#3&quot;, CREATE] 186 [Email, &quot;B&quot;, UPDATE] 187 </code></pre><p>Because <code>A</code> is back to its original state, we’ve eliminated it entirely from the log, and do not need to send anything to the server. Because <code>#2</code> was a create in the log, we do not move it when we renamed it at the end, which is good because otherwise the other changes in the log would both fail as they depend on it. Despite making two changes to <code>B</code>, we only have to send a single update to the server for it.</p><h2 id="conflicts" tabindex="-1">Conflicts</h2><p>Suppose you have a shared contact, let’s call him Joe Bloggs. While offline you edit to add his phone number. Meanwhile, a colleague updates his email address. This means when your client comes back online and synchronises the changes, the object it is updating has already changed. This is called a conflict.</p><p>For the data types we have to handle, we believe automatic resolution (rather than presenting the conflict to the user and asking them to choose what should happen) is the right way to go. We follow these simple rules:</p><ul> <li>Last write wins.</li> <li>All updates are patches.</li> </ul><p>This means if the same object is updated by two different people, whichever client writes second will overwrite the data of the one that wrote first. (The first client will then sync this change back so you get a consistent state.) However, since all updates are patches, it will merge the changes unless they apply to the same property on the object. So in the case above, although there were two writes to the same contact, they were updating different properties. One user was updating the “<a href="https://www.rfc-editor.org/rfc/rfc9553.html#name-emails" target="_blank" rel="noopener">emails</a>”, the other the “<a href="https://www.rfc-editor.org/rfc/rfc9553.html#name-phones" target="_blank" rel="noopener">phones</a>”. So in this case, both changes would be preserved.</p><h2 id="next-up-mail-storage" tabindex="-1">Next up, mail storage</h2><p>In this post we looked at how we store changes you make offline so we can accurately and efficiently sync them back to the server when you come online. Like our discussion of data storage yesterday, everything here applies generically to all data types.</p><p>Tomorrow, we’ll discuss <a href="/blog/offline-mail-storage/">why email is special</a>, and what else we do to make this super fast in our offline store.</p></content> 188 </entry><entry> 189 <title>Dec 17: Building offline: general architecture</title> 190 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/offline-architecture/' /> 191 <id>https://www.fastmail.com/blog/offline-architecture/</id> 192 <updated>2024-12-17T00:00:01Z</updated><author> 193 <name>Neil Jenkins</name> 194 </author><content xml:lang='en' type='html'><p>This is the seventeenth post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/offline-in-beta/">Dec 16: Offline support now in public beta</a>. The next post is <a href="/blog/offline-sync/">Dec 18: Building offline: syncing changes back to the server</a>.</p><p>Yesterday <a href="/blog/offline-in-beta/">we announced full offline support for Fastmail</a> is now available in public beta, both in our app and on the web. Today, and over the next few days, I’ll dive into some of the technical aspects about how we’re making this work.</p><p>Making Fastmail work offline has been our most popular feature request for some time (ever since <a href="/blog/more-swipe-options-on-mobile-let-you-work-faster/">we added support for custom swipe actions</a>, our previous top request!), and we wanted to make sure our support was done <em>right</em>. It should just work, seamlessly, and as far as possible you should be able to do everything you can do online. Open your calendar and update an event, perhaps inviting someone using autocomplete from your contacts. Search your mail and triage it. Write a new note. Add a memo. We want it all to <em>just work</em>.</p><p>When you come online it should seamlessly sync these changes back to the server, and fetch any new mail and other changes.</p><h2 id="general-architecture" tabindex="-1">General architecture</h2><p>The Fastmail app is very cleanly separated from our server, which was a huge benefit when adding offline support. All data in the app is loaded via a <a href="https://jmap.io/" target="_blank" rel="noopener">JMAP</a> API, with all UI rendering and routing happening in the app. This means the data flow looks a bit like this:</p><pre><code>[App] ← JMAP → [Server] 195 </code></pre><p>This gave us a really well defined boundary on which to build the offline support. If we built something that could handle and respond to the JMAP requests directly on your device then the app would work offline, and we wouldn’t really have to change anything else in it. So the architecture we came up with looks like this:</p><pre><code>[App] ← JMAP → [Caching layer] ← JMAP → [Server] 196 </code></pre><p>Looking at our new diagram, we can see that we are building two things:</p><ol> <li>A JMAP server that can understand the requests the client makes, fetch and write the data from/to a local store, and return a JMAP response.</li> <li>A JMAP client that can fetch data it doesn’t have from the server, and write back changes made while we were offline.</li> </ol><p>This is actually quite a lot harder than just building a JMAP server! We will often only have partial information, and we have to transparently pass through requests for data we don’t have to the server, and gracefully handle a fallback if we are offline.</p><p>The core function our caching layer needs to perform is handling a JMAP request from the client. Each JMAP request is a sequence of method calls. Once we have locally cached data, we may be able to handle the method call entirely locally, but we may always run into one or more that we can’t.</p><p>For performance, we want to batch our method calls into a single database transaction where possible. But we can’t hold open a transaction over a network request, so we divide up the execution into phases.</p><p>First, we attempt to execute the method calls locally. If all complete, we’re done! If any require us to fallback to the server, we stop and send it everything from that point on, as we want to avoid making multiple HTTP requests, which would be slow.</p><p>If the request completed successfully, we process the responses to save any new data into our local datastore. If the request failed, we call the offline fallback methods, which may be able to still return a response to the UI.</p><h2 id="a-separate-thread-for-the-caching-layer" tabindex="-1">A separate thread for the caching layer</h2><p>The caching layer runs on your device, just like the rest of the app. Because JMAP requests are already asynchronous network calls from the UI, we can easily run the caching layer in a separate OS thread so it never blocks the UI thread, which could cause <a href="https://en.wiktionary.org/wiki/jank" target="_blank" rel="noopener">jank</a>.</p><p>Our app is built using web technology, which allows us, as a small company, to build an app that runs everywhere our users are, with a single code base and feature parity across all platforms. Separate threads are represented as <a href="https://developer.mozilla.org/en-US/docs/Web/API/Web_Workers_API" target="_blank" rel="noopener">workers</a> in the web API. There are three types of worker:</p><ul> <li><a href="https://developer.mozilla.org/en-US/docs/Web/API/Worker" target="_blank" rel="noopener">Dedicated worker</a> — this is a worker that is tied to a particular window or tab in your browser. If you have Fastmail open in multiple tabs, each would have to create its own worker.</li> <li><a href="https://developer.mozilla.org/en-US/docs/Web/API/SharedWorker" target="_blank" rel="noopener">Shared worker</a> — this is a worker that’s shared between tabs or windows, so no matter how many you have there’s only a single instance of this worker.</li> <li><a href="https://developer.mozilla.org/en-US/docs/Web/API/Service_Worker_API" target="_blank" rel="noopener">Service worker</a> — this is a special type of shared worker that can intercept network requests and change their response.</li> </ul><p>At first glance, a service worker seems the place to handle all of this, and this is what we tried first. However, we soon switched over to using a shared worker instead:</p><ul> <li>The service worker is designed to be short lived and only spun up when needed, but we wanted to hold open a persistent <a href="https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_events" target="_blank" rel="noopener">EventSource push connection</a> while the app is open, for instant updates. The shared worker is a better conceptual fit for this.</li> <li>We don’t need to intercept network requests, as we can just pass the JMAP request object directly to the worker using the <a href="https://developer.mozilla.org/en-US/docs/Web/API/Worker/postMessage" target="_blank" rel="noopener">postMessage</a> API, avoiding some serialisation overhead.</li> <li>We ran into a bug in iOS where network requests would sometimes not be intercepted even though the service worker was registered when the app was running in the background. This meant we <em>had</em> to pass the request directly to the worker anyway instead of intercepting it at the network level to ensure we didn’t hit this bug.</li> </ul><p>We wanted to use a shared worker rather than a dedicated worker for more efficiency when there were multiple tabs open — we can avoid some contention and locking issues, and ensure we have a single push connection open to the server.</p><h2 id="storing-data-indexed-db" tabindex="-1">Storing data: IndexedDB</h2><p>The web API for storing large volumes of structured data is <a href="https://developer.mozilla.org/en-US/docs/Web/API/IndexedDB_API" target="_blank" rel="noopener">IndexedDB</a>. This lets you create multiple <a href="https://developer.mozilla.org/en-US/docs/Web/API/IDBObjectStore" target="_blank" rel="noopener">object stores</a> (the equivalent of SQL tables), which offer simple key-value storage with ordered keys. Indexes can be automatically built based on properties in the object being stored. Transactions ensure data consistency.</p><p>The IndexedDB API was unfortunately designed just before promises became ubiquitous in the web world. This means just fetching a record from a store requires code a bit like this:</p><pre class="language-javascript"><code class="language-javascript"><span class="keyword token">const</span> request <span class="operator token">=</span> store<span class="punctuation token">.</span><span class="function token">get</span><span class="punctuation token">(</span>id<span class="punctuation token">)</span><span class="punctuation token">;</span> 197 request<span class="punctuation token">.</span><span class="function function-variable token">onerror</span> <span class="operator token">=</span> <span class="punctuation token">(</span><span class="punctuation token">)</span> <span class="operator token">=></span> <span class="function token">callErrorHandler</span><span class="punctuation token">(</span><span class="punctuation token">)</span><span class="punctuation token">;</span> 198 request<span class="punctuation token">.</span><span class="function function-variable token">onsuccess</span> <span class="operator token">=</span> <span class="punctuation token">(</span><span class="punctuation token">)</span> <span class="operator token">=></span> <span class="punctuation token">{</span> 199 <span class="keyword token">const</span> record <span class="operator token">=</span> request<span class="punctuation token">.</span>result<span class="punctuation token">;</span> 200 <span class="comment token">// Do something</span> 201 <span class="punctuation token">}</span><span class="punctuation token">;</span></code></pre><p>This is clunky and becomes hard to read and follow. However, we wrote one tiny little wrapper function that converts it into a promise-based API:</p><pre class="language-javascript"><code class="language-javascript"><span class="keyword token">const</span> <span class="function function-variable token">_</span> <span class="operator token">=</span> <span class="punctuation token">(</span><span class="parameter token">request</span><span class="punctuation token">)</span> <span class="operator token">=></span> 202 <span class="keyword token">new</span> <span class="class-name token">Promise</span><span class="punctuation token">(</span><span class="punctuation token">(</span><span class="parameter token">resolve<span class="punctuation token">,</span> reject</span><span class="punctuation token">)</span> <span class="operator token">=></span> <span class="punctuation token">{</span> 203 request<span class="punctuation token">.</span><span class="function function-variable token">onsuccess</span> <span class="operator token">=</span> <span class="punctuation token">(</span><span class="punctuation token">)</span> <span class="operator token">=></span> <span class="function token">resolve</span><span class="punctuation token">(</span>request<span class="punctuation token">.</span>result<span class="punctuation token">)</span><span class="punctuation token">;</span> 204 request<span class="punctuation token">.</span><span class="function function-variable token">onerror</span> <span class="operator token">=</span> <span class="punctuation token">(</span><span class="punctuation token">)</span> <span class="operator token">=></span> <span class="function token">reject</span><span class="punctuation token">(</span>request<span class="punctuation token">.</span>error<span class="punctuation token">)</span><span class="punctuation token">;</span> 205 <span class="punctuation token">}</span><span class="punctuation token">)</span><span class="punctuation token">;</span></code></pre><p>Using this function, we can rewrite the above fetch like this:</p><pre class="language-javascript"><code class="language-javascript"><span class="keyword token">const</span> record <span class="operator token">=</span> <span class="keyword token">await</span> <span class="function token">_</span><span class="punctuation token">(</span>store<span class="punctuation token">.</span><span class="function token">get</span><span class="punctuation token">(</span>id<span class="punctuation token">)</span><span class="punctuation token">)</span><span class="punctuation token">;</span> 206 <span class="comment token">// Do something</span></code></pre><p>(Note, if the fetch has an error this will result in an exception being thrown, which is generally handled at a higher layer, so avoids that cluttering our code here at all!)</p><p>With this simple addition, I found the IndexedDB API consistent and easy to work with.</p><h2 id="the-standard-object-store-structure" tabindex="-1">The standard object store structure</h2><p>The consistency of JMAP means we can write one generic implementation and then use it to provide offline support for all our data types. For each data type (such as Calendar, Email, Contact, etc.) we create an object store to store the instances of that type.</p><p>When we create the object store we also store a single metadata object in it, using a special key (a zero-byte <a href="https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/ArrayBuffer" target="_blank" rel="noopener">ArrayBuffer</a>). This stores some important bookkeeping information, in particular the following properties:</p><pre class="language-javascript"><code class="language-javascript"><span class="comment token">// Do we have the full set of data from the server for this data</span> 207 <span class="comment token">// type?</span> 208 <span class="literal-property property token">hasAllRecords</span><span class="operator token">:</span> <span class="boolean token">false</span><span class="punctuation token">,</span> 209 210 <span class="comment token">// State string representing the current server state we have</span> 211 <span class="comment token">// synced with. The store may also contain newer information, but</span> 212 <span class="comment token">// that's ok as it will still get to the correct state when we</span> 213 <span class="comment token">// update from the old state.</span> 214 <span class="literal-property property token">serverState</span><span class="operator token">:</span> <span class="string token">''</span><span class="punctuation token">,</span> 215 216 <span class="comment token">// This is the highest modseq of a record in the store.</span> 217 <span class="literal-property property token">lastModSeq</span><span class="operator token">:</span> <span class="number token">0</span><span class="punctuation token">,</span> 218 219 <span class="comment token">// This is the highest modseq of a record that was destroyed that's</span> 220 <span class="comment token">// now been removed entirely from the store; we can only calculate</span> 221 <span class="comment token">// changes accurately from this point on.</span> 222 <span class="literal-property property token">highestPurgedModSeq</span><span class="operator token">:</span> <span class="number token">0</span><span class="punctuation token">,</span> 223 224 <span class="comment token">// This is the number of records currently marked destroyed in the</span> 225 <span class="comment token">// store. We keep them there so we can calculate changes. Once we</span> 226 <span class="comment token">// cross a threshold, we'll clean up old ones.</span> 227 <span class="literal-property property token">numDestroyed</span><span class="operator token">:</span> <span class="number token">0</span><span class="punctuation token">,</span></code></pre><p>A key concept here is <em>modseq</em>, which stands for “modification sequence”. It’s a counter we keep per account, per data type. Every time we make a change to a record in our local store we bump the sequence number and assign that as the new “updated” modseq for that record. We also store a “created” modseq on each record, which is the same as the “updated” modseq when the record is first created. These simple bookkeeping properties allow us to efficiently calculate changes, as needed for <a href="https://www.rfc-editor.org/rfc/rfc8620.html#section-5.2" target="_blank" rel="noopener">the JMAP “/changes” method</a>.</p><p>Aside from the metadata object, every other entry in the object store is a record — an instance of that data type.</p><p>The key for each record is <code>[account id, id]</code>, because some data types exist in multiple accounts (e.g. shared contacts and your personal contacts) and ids are only unique within an account. As far as I could see from inspecting the source, string keys are stored as <a href="https://en.wikipedia.org/wiki/UTF-16" target="_blank" rel="noopener">UTF-16</a> in all major IndexedDB implementations. This is a fairly inefficient encoding, especially as we know JMAP ids can only use the <a href="https://datatracker.ietf.org/doc/html/rfc4648#section-5" target="_blank" rel="noopener">base64url characters</a>, so for efficiency we encode this data into an <a href="https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/ArrayBuffer" target="_blank" rel="noopener">ArrayBuffer</a>, making use of this fact.</p><p>The value associated with the key is the record itself — an object representing an instance of that data type, as fetched from the server. In addition, we add a few bookkeeping properties, as discussed above:</p><ul> <li>The created modseq</li> <li>The updated modseq</li> <li>Is the record destroyed?</li> </ul><p>Each object store has an index built automatically based on the updated modseq of the records.</p><p>The modseq is used as the “state” string over JMAP. When asked for what’s changed since a particular state, we know:</p><ul> <li>Only records with a higher modseq have changed (which we can efficiently get from the index).</li> <li>If the record’s <code>created</code> modseq is higher than <code>lastModSeq</code>, it’s new. Otherwise it’s been updated or destroyed (depending on whether the record is now destroyed). If it’s new and also destroyed, we can ignore it entirely.</li> </ul><h2 id="next-up-local-changes" tabindex="-1">Next up, local changes</h2><p>In this post we looked at the basic overview of how our offline caching layer fits into our app, and the way it stores data to efficiently respond to JMAP requests. Tomorrow, we’ll dive into <a href="/blog/offline-sync/">how it keeps track of changes the user makes while offline</a>, so it can reconcile this with the server.</p></content> 228 </entry><entry> 229 <title>Dec 16: Offline support now in public beta</title> 230 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/offline-in-beta/' /> 231 <id>https://www.fastmail.com/blog/offline-in-beta/</id> 232 <updated>2024-12-16T00:00:01Z</updated><author> 233 <name>Neil Jenkins</name> 234 </author><content xml:lang='en' type='html'><p>This is the sixteenth post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/platform-team-working-agreement/">Dec 15: Platform Team working agreement</a>. The next post is <a href="/blog/offline-architecture/">Dec 17: Building offline: general architecture</a>.</p><p>On an aeroplane, or down a tunnel. Roaming in a foreign land, or just out for a walk in the country. There are still times when we aren’t connected to the internet. As a strong supporter of open standards, you’ve always been able to use Fastmail with any email app you like, many of which work offline. However, many of our customers prefer the Fastmail app (we certainly do!), and offline support there has been our most popular request for some time. As Bron <a href="/blog/2023-advent-post-bron-gondwana/">foreshadowed last year</a>, we’ve been working hard on this big project for some time, and and we’re very pleased to announce it’s now ready for public beta testing.</p><h2 id="what-s-a-beta" tabindex="-1">What’s a beta?</h2><p>Fastmail runs a beta version of our app to allow our users that like to live on the cutting edge to get early access to new features before they’re finished, and send us feedback while we’re still working on them. Things may be broken occasionally, but generally it’s pretty stable.</p><h2 id="how-do-i-get-the-beta" tabindex="-1">How do I get the beta?</h2><p>If you are using our iOS or Android apps:</p><ol> <li>Check for updates in the app store — make sure you have the latest version.</li> <li>Go to Settings → Device settings → Show Advanced Settings, and change the “server backend” to “Beta”.</li> </ol><p>If you are using Fastmail on the web, just log in at <a href="https://betaapp.fastmail.com" target="_blank" rel="noopener">betaapp.fastmail.com</a>.</p><p>In either case, once on beta you will find a new settings screen — <a href="https://betaapp.fastmail.com/settings/offline" target="_blank" rel="noopener">Settings → Offline</a> — where you can toggle on offline support. You must have ticked “Keep me logged in” when you logged in to enable offline support.</p><h2 id="what-can-i-do-offline" tabindex="-1">What can I do offline?</h2><p>Almost everything! You can read mail, reply, view and edit your contacts/calendar, change settings, etc.…</p><p>Probably easier is to describe what <em>won’t</em> work offline:</p><ul> <li>Mail search will not look inside attachments, and will give slightly different results to when online. If you don’t choose to make every message available offline, it won’t be able to match against content it hasn’t downloaded!</li> <li>Snoozed messages will not move back to the inbox while offline.</li> <li>Calendar reminders will not show a notification.</li> <li>You can’t delete attachments.</li> <li>You can’t attach files to a message you are composing.</li> <li>You can’t add or change users or domains, change your plan or update your billing details, or change your security settings.</li> </ul><h2 id="how-do-i-send-feedback" tabindex="-1">How do I send feedback?</h2><p>Found a bug? Got an idea? Something you like, or don’t like? <a href="/support/">Let us know</a>! Our support team will pass all feedback on to the product team. We promise we read and carefully consider it all.</p><h2 id="what-s-the-tech-behind-your-offline-support" tabindex="-1">What’s the tech behind your offline support?</h2><p>Interested in the technical details? Over the next few days I’m going to dive into how we built offline support into our app, starting with <a href="/blog/offline-architecture/">the general architecture</a>.</p></content> 235 </entry><entry> 236 <title>Dec 15: Platform Team working agreement</title> 237 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/platform-team-working-agreement/' /> 238 <id>https://www.fastmail.com/blog/platform-team-working-agreement/</id> 239 <updated>2024-12-15T00:00:01Z</updated><author> 240 <name>Bron Gondwana</name> 241 </author><content xml:lang='en' type='html'><p>This is the fifteenth post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/on-call-systems/">Dec 14: On-call systems</a>. The next post is <a href="/blog/offline-in-beta/">Dec 16: Offline support now in public beta</a>.</p><p>Another Sunday, another document from our internal collection!</p><p>Over the past couple of years, we’ve been taking the bits that we like from scrum, agile, and all the other buzzword methodologies that are out there. One of the best things that all the good frameworks to is give a space for setting behaviours for ourselves and each other within a team — so you know what to expect from your colleagues, and also what they expect of you!</p><p>I posted about our Mission Statement and Guiding Principles already, now we narrow down to look at a single team. Our platform team maintains the infrastructure that the Fastmail product runs on top of.</p><p>We started from here:</p><h2 id="principles-and-duties" tabindex="-1">Principles and duties</h2><p><strong>We are lifeguards:</strong></p><ul> <li>We watch out for upcoming risks and for systems in distress.</li> <li>Part of our day is keeping a watchful eye and scanning for strangeness and hints of something going down.</li> </ul><p><strong>We are first-aiders:</strong></p><ul> <li>When something goes wrong, our first principle is to keep the patient alive</li> <li>For Fastmail: the patient is the data stored on our systems; the data we store on the behalf of others. This is a higher priority than uptime.</li> </ul><p><strong>We are paramedics:</strong></p><ul> <li>We do more than just keep the patient alive! We perform surgery and repairs to the level of our ability.</li> <li>For more complex things, we stabilise and bring things to the specialists in that system to make more permanent repairs.</li> <li>We get the service back up for as many as possible, as quickly as possible, without compromising data.</li> </ul><p><strong>We are specialists:</strong></p><ul> <li>For the basics of first aid and early response, we all need to become experts - and we all need to be vigilant.</li> <li>But; we all have our areas of specialty, and it’s OK to lean into that!</li> </ul><p><strong>The duties of a platformer:</strong></p><ul> <li>Monitoring - looking at the alert emails, the metrics, etc and making sure things are nominal.</li> <li>First-aid first - if the platform is in crisis, drop everything and fix it. This is the oncall duty, we all do this.</li> <li>Janitorial - mop the floors, oil the joints, replace worn parts. We all do this, on rotation and as we see things. We keep the ship ship-shape.</li> <li>Initiatives - look into possible improvements, experiment with possibilities, and build the future. Everyone should have an initiative they are leading or collaborating on. These make progress in the time we’re not doing first-aid or janitorial work.</li> </ul><p>A platform team shouldn’t be particularly busy. The first-aid and janitorial work should not be a large chunk of your day most of the time.</p><p>Everyone on the platform team is an experienced and senior Platform Engineer, though most of the current team are quite new to Fastmail, which is why I’m also embedded with them as the crusty old expert who knows where many of the bodies are buried!</p><p>Since the team is split between USA and Australia; we only spend an hour of dedicated synchronous time per week. So we brought everyone together in Melbourne in November, and came up with the following schedule and working agreement:</p><p>Over each 4 week period, this includes two set of task list gardening, two technical deep-dive sessions, a retrospective where we can change how we operate as a team! One member of the team runs all the ceremonies for a 4 week period, and hands off after the retrospective.</p><h2 id="working-agreement" tabindex="-1">Working Agreement</h2><p>We value action, forward progress and take ownership of tasks</p><p>We learn from each others successes and mistakes, and use good processes to protect ourselves from human error, always asking how we can be better.</p><p>We give and seek feedback, and encourage asking questions (no dumb questions)</p><p>We generate discussion and directional buy in, and make time per site for mobbing/pairing and getting clarity on solutions.</p><p>We celebrate our successes and show them off to the rest of the company.</p><p><strong>To resolve conflict we:</strong></p><ul> <li>Check with the customer (the Fastmail product team) - what do they need?</li> <li>Check with an area expert</li> <li>Use experiments</li> </ul><p><strong>To agree things we:</strong></p><ul> <li>Make solo calls if comfortable; or</li> <li>Bring it ‘To Discuss’ - the fortnightly deep dive meeting</li> </ul></content> 242 </entry><entry> 243 <title>Dec 14: On-call systems</title> 244 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/on-call-systems/' /> 245 <id>https://www.fastmail.com/blog/on-call-systems/</id> 246 <updated>2024-12-14T00:00:02Z</updated><author> 247 <name>Luke Erlacher</name> 248 </author><content xml:lang='en' type='html'><p>This is the fourteenth post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/moving-fastmail-dns-to-knot/">Dec 13: It’s knot DNS. There’s no way it’s DNS. It is DNS!</a> The next post is <a href="/blog/platform-team-working-agreement/">Dec 15: Platform Team working agreement</a>.</p><p>After the scary interlude for Friday the 13th, here’s a follow-up to our blog about support and on-call.</p><p>At Fastmail, we own and operate most of our stack, from networking gear and servers in the hosted datacenters we use, server OS setup, internal services, backups and redundancy, through to user-facing services and mail flow.</p><p>In the last post, we introduced our organizational model of how our support and on-call team manage on-call and incidents. In this blog post, we get a little more technical into the systems we use to help us do this.</p><h2 id="how-incidents-are-raised" tabindex="-1">How Incidents are raised</h2><p>We source alerts from our service stack through a number of mechanisms:</p><ul> <li>Prometheus alerts for infrastructure and service metrics such as disk usage and replication delay</li> <li>Logwatchers that alert on things like Cyrus errors</li> <li>Cron scripts that run regular tests and alert on failures</li> <li>Alerts from external monitoring</li> </ul><p>We’re currently experimenting with integrating all of these sources into a single observability stack (I’m a big fan of the unified observability model) - maybe you will read about that in next year’s advent blog post series!</p><p>But right now, all of these alerts go straight to Pagerduty to create an incident there.</p><p>Our support team and any staff member is empowered to raise an incident at any time when they observe issues that indicate an infrastructure or service failure. This is done via slack pings during working hours, and otherwise by paging through our slackops bot.</p><h2 id="pagerduty-alerts" tabindex="-1">Pagerduty alerts</h2><p>Our Pagerduty on-call rotation and escalation process ensure that incidents are always promptly responded to.</p><p>The primary on-call engineer is also responsible for Business As Usual (BAU) tasks such as reviewing non-critical errors and escalations. When incidents have to be escalated past the primary on-call engineer, they go first to the team lead. This allows the on-call engineers that are off rotation to focus on project work during the day, and decompress outside of work hours without worrying about on-call.</p><p>In order to minimize the attention load of on-call alerts on the platform engineers, engineers can override the on-call schedule to take on-call when they do large deploys that result in alerts.</p><p>We have recently created a dedicated incident discussion channel. Previously, incident communication would take place either in the alerts feed channel, the “ops” channel where we post updates for visibility into changes and deploys, or in a team channel for the team principally responding to the incident.</p><p>A dedicated incident discussion channel removes incident discussion noise from other channels. We decided against going with the “create a dedicated channel for every incident” as we think this would create too much churn and reduce visibility.</p><h2 id="outage-notifications" tabindex="-1">Outage notifications</h2><p>When our service is experiencing issues or outages, we want our users to know as soon as possible, both for their benefit and our support team. As we have support staff on shift 24/7, they handle outage notifications via our status page on https://fastmailstatus.com/, Zendesk banners, and posting on our social channels. During incidents, they act as the communication conduit to our users - ascertaining the extent of outages, ensuring timely updates, and checking that when services are restored this is reflected in user experience.</p><h2 id="pagerduty-nitty-gritty" tabindex="-1">Pagerduty nitty-gritty</h2><p>Some of the things we do to automate and simplify our life as on-call engineers butt up against the limits of Pagerduty.</p><p>For example, it would be nice to have an empty escalation layer as the first layer that during normal times falls through to the next layer instantly. Then people can add themselves in this layer to override the on-call escalation chain for doing deploys. However Pagerduty does not allow this so we have to override the schedules for primary on-call instead.</p><p>We also have regular unavailabilities during the week for on-call engineers for scheduled events such as gym training or dance classes. For this, we want to add exclusions to the schedule so that the alerts go to the fallback on-call immediately.</p><p>These unavailabilities are different for every individual engineer. Pagerduty doesn’t allow to model this - on-call schedules can have almost arbitrary scheduling through weekdays, but this is per schedule and not per on-call user. So we have to split the schedule and make a separate schedule for every engineer. However, we then can no longer make a rotating schedule - inside one schedule, a user can’t be on for a week and then off for a week - unless that off week is taken by another user.</p><p>To work around this we have created (and paid for) a “Blank” dummy user that has no contact methods and is only there to fill the off week for a user.</p><p><a href="/assets/blog/2024-12-14-oncall-systems/pagerduty_screenshot_3x.png" target="_blank"><picture><source type="image/webp" srcset="/assets/images/pagerduty_screenshot_3x-4tkoE_Vw31-375.webp 375w, /assets/images/pagerduty_screenshot_3x-4tkoE_Vw31-750.webp 750w, /assets/images/pagerduty_screenshot_3x-4tkoE_Vw31-1500.webp 1500w" sizes="(max-width: 425px) 375px, 750px"><img alt="Pagerduty schedule" loading="lazy" decoding="async" src="/assets/images/pagerduty_screenshot_3x-4tkoE_Vw31-375.png" width="1500" height="514" srcset="/assets/images/pagerduty_screenshot_3x-4tkoE_Vw31-375.png 375w, /assets/images/pagerduty_screenshot_3x-4tkoE_Vw31-750.png 750w, /assets/images/pagerduty_screenshot_3x-4tkoE_Vw31-1500.png 1500w" sizes="(max-width: 425px) 375px, 750px"></picture></a></p><p>As you can see in the picture, there are 3 layers in the schedule: The first layer is a base layer that is a weekly rotation between the fallback engineer and the Blank user.</p><p>The second layer has the primary oncall engineer’s shift for the first half of the day, and the third layer has the second half of the day. This is also in a weekly rotation between the oncall engineer and the Blank user.</p><p>To understand what’s going on here, imagine that the primary oncall engineer needs to be offline every monday from 4PM to 6PM. For that, we put an 8AM - 4PM shift time for every monday in the second layer, and 6PM to 10PM in the third layer. During the 4PM to 6PM time, there is no on-call user in the second or third layer, so it falls back to the base layer and so the fallback engineer will be active.</p><p>Finally, we make a separate schedule for every on-call engineer following the same pattern, but offset by one week.</p><h2 id="what-works-well" tabindex="-1">What works well</h2><p>I have done overnight on-call at previous companies and that is a lot more stressful. I definitely think being able to split on-call is a lot better for engineers’ quality of life!</p><p>We have ops tools to do things like log searching, show replication stats, grafana dashboards, and prometheus alerts that give us quick insights into the state of our infrastructure and services to pin down the root cause of an incident quickly. These could be more comprehensive and better organized but they work well.</p><p>We have runbooks for some, but not all, alerts. Where they exist they are quite good and comprehensive and we frequently review and update them after incidents.</p><h2 id="what-could-work-better" tabindex="-1">What could work better</h2><p>Most of our incidents auto-resolve when the service / monitor returns to nominal service. This is good! We are lucky to have very few flappy alerts that bounce up and down.</p><p>But a handful of alerts do not auto-resolve because the alerting mechanism is not stateful. This is confusing for engineers, and it is an extra annoyance to clean up alerts after an incident.</p><p>We don’t currently have good categorization and prioritization of alerts. The only way to know whether an alert is important or not is to know from experience, or asking someone with experience. We need to spend more time and be more ruthless in weeding out low-quality alerts!</p></content> 249 </entry><entry> 250 <title>Dec 13: It’s knot DNS. There’s no way it’s DNS. It is DNS!</title> 251 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/moving-fastmail-dns-to-knot/' /> 252 <id>https://www.fastmail.com/blog/moving-fastmail-dns-to-knot/</id> 253 <updated>2024-12-13T00:00:01Z</updated><author> 254 <name>Rob Mueller</name> 255 </author><content xml:lang='en' type='html'><p>This is the thirteenth post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/following-the-sun/">Dec 12: Following the Sun</a>. The next post is <a href="/blog/on-call-systems/">Dec 14: On-call systems</a>.</p><p>Ten years ago on December 13, 2014, I talked about how Fastmail had moved its DNS system <a href="/blog/fastmail-dns-hosting/">from TinyDNS to PowerDNS</a>.</p><p>In 2023, we made another big move, switching all our DNS serving from PowerDNS to <a href="https://www.knot-dns.cz/" target="_blank" rel="noopener">Knot DNS</a>. This turned out to be a fairly large change and we ended up going down a couple of different paths before landing on a solid implementation.</p><h2 id="where-were-we-up-to-again" tabindex="-1">Where were we up to again?</h2><p>As a reminder where we <a href="/blog/fastmail-dns-hosting/">left off in 2014</a>, we’d moved DNS serving for our <code>ns[12].messagingengine.com</code> DNS servers from static TinyDNS files to using PowerDNS with its <a href="https://doc.powerdns.com/authoritative/backends/pipe.html" target="_blank" rel="noopener">pipe backend</a> to generate content dynamically on each DNS request via internal logic in code. Modulo some caching, this removed the latency from when you made an update to your domain’s DNS in our UI to when those changes became visible at our DNS servers.</p><p>If you own your own domain and host it at Fastmail, it’s easy to customise the DNS for it. Just go to <a href="https://app.fastmail.com/settings/domains" target="_blank" rel="noopener">Settings -&gt; Domains</a> and click <strong>Edit</strong> next to the domain.</p><p><picture><source type="image/webp" srcset="/assets/images/customise-dns-1-F6B9iDf5KA-375.webp 375w, /assets/images/customise-dns-1-F6B9iDf5KA-750.webp 750w, /assets/images/customise-dns-1-F6B9iDf5KA-1024.webp 1024w" sizes="(max-width: 425px) 375px, 750px"><img alt="Screenshot of domain settings" loading="lazy" decoding="async" src="/assets/images/customise-dns-1-F6B9iDf5KA-375.png" width="1024" height="100" srcset="/assets/images/customise-dns-1-F6B9iDf5KA-375.png 375w, /assets/images/customise-dns-1-F6B9iDf5KA-750.png 750w, /assets/images/customise-dns-1-F6B9iDf5KA-1024.png 1024w" sizes="(max-width: 425px) 375px, 750px"></picture></p><p>Then <strong>Customise DNS</strong>.</p><p><picture><source type="image/webp" srcset="/assets/images/customise-dns-2-q21fuWTaz8-375.webp 375w, /assets/images/customise-dns-2-q21fuWTaz8-750.webp 750w, /assets/images/customise-dns-2-q21fuWTaz8-1024.webp 1024w" sizes="(max-width: 425px) 375px, 750px"><img alt="Screenshot of Customise DNS link" loading="lazy" decoding="async" src="/assets/images/customise-dns-2-q21fuWTaz8-375.png" width="1024" height="80" srcset="/assets/images/customise-dns-2-q21fuWTaz8-375.png 375w, /assets/images/customise-dns-2-q21fuWTaz8-750.png 750w, /assets/images/customise-dns-2-q21fuWTaz8-1024.png 1024w" sizes="(max-width: 425px) 375px, 750px"></picture></p><p>By default we generate a number of records to make using your domain with email easy. We recommend leaving these as is, but you have full control and it’s easy to add or change additional records.</p><p><picture><source type="image/webp" srcset="/assets/images/customise-dns-3-_fxZZelxrZ-375.webp 375w, /assets/images/customise-dns-3-_fxZZelxrZ-750.webp 750w, /assets/images/customise-dns-3-_fxZZelxrZ-1024.webp 1024w" sizes="(max-width: 425px) 375px, 750px"><img alt="Screenshot of custom DNS page" loading="lazy" decoding="async" src="/assets/images/customise-dns-3-_fxZZelxrZ-375.png" width="1024" height="270" srcset="/assets/images/customise-dns-3-_fxZZelxrZ-375.png 375w, /assets/images/customise-dns-3-_fxZZelxrZ-750.png 750w, /assets/images/customise-dns-3-_fxZZelxrZ-1024.png 1024w" sizes="(max-width: 425px) 375px, 750px"></picture></p><pre><code># dig +short a.b.c.uberengineer.com TXT 256 &quot;txt for a.b.c&quot; 257 </code></pre><p>While this solved the original latency problem we had, it introduced a few others.</p><p>The pipe backend was computationally considerably more expensive. Under normal DNS load this wasn’t a problem and the system was scaled appropriately to handle it just fine. However if we got hit with excessive load mostly due to some form of <a href="https://en.wikipedia.org/wiki/Denial-of-service_attack#Distributed_DoS_attack" target="_blank" rel="noopener">DDoS</a> attack, the system could easily come under strain and start to fail. When this happens people can’t access our website reliably, or our IMAP/POP/SMTP servers, and worst of all other sites might have problems working out which servers to deliver email for Fastmail customers to.</p><p>To protect our servers from DDoS attacks, we had put them behind <a href="https://developers.cloudflare.com/dns/dns-firewall/" target="_blank" rel="noopener">Cloudflare’s DNS firewall</a> product, which is a global distributed DNS cache.</p><p>The Cloudflare DNS firewall product works great at edge caching if there’s a flood of DNS queries to a particular domain (or small set of) domain names and solved most of the DDoS flooding issues we saw. However we experienced cases where we were flooded with DNS queries of the form <code>$randomdomain.fastmail.com</code> (called a <a href="https://developers.cloudflare.com/dns/dns-firewall/random-prefix-attacks/about/" target="_blank" rel="noopener">pseudo random prefix attack</a> or random subdomain attack). If every DNS query is to a random different sub-domain, then Cloudflare has to pass all those queries straight through to us as it has nothing cached. In theory Cloudflare say they can mitigate this. Unfortunately we felt their mitigation didn’t work particularly well and still resulted in a significant overload of incoming DNS queries.</p><p>Again, this flood of queries could cause an overload of the PowerDNS pipe backend processes which caused visible DNS downtime. We couldn’t find any sensible tuning that would allow PowerDNS and our pipe backend to correctly operate under one of these sustained floods.</p><h2 id="mitigating-random-prefix-attacks" tabindex="-1">Mitigating random prefix attacks</h2><p>At the time we needed a quick solution to this problem. So we ended up splitting our DNS in two. We noted that basically all the random prefix attacks were against our system domains like fastmail.com and not user domains. Since DNS for our system domains doesn’t change much at all we basically backtracked and put our system domains onto separate nameservers that ran TinyDNS using a mostly static database of DNS records. Although this felt hacky, it worked. TinyDNS was able to absorb the higher query load in these attack situations quite well.</p><p>This however complicated our DNS setup even more. We also knew if a user domain experienced one of these attacks, it wasn’t using the TinyDNS system. Obviously it was possible an attack could take out not just the one user domain, but <em>all</em> user domains using our DNS if it overloaded the PowerDNS server. We wanted a simpler, better, and more permanent solution.</p><h2 id="replacing-all-dns-with-knot-dns-server" tabindex="-1">Replacing all DNS with Knot DNS server</h2><p>So the main observation about DNS is that in general it doesn’t actually change that often. So what we wanted was a solution where each of the 100,000’s of domains in our system represents a DNS zone that can be individually built into a static database, and each zone can be added/updated/deleted separately without having to reload/rebuild all zones.</p><p>We looked around at a few servers and went with <a href="https://www.knot-dns.cz/" target="_blank" rel="noopener">Knot DNS</a>.</p><ul> <li>It looks like a “standard” DNS server with zone files, replication, etc, so it’s easier for new staff to understand</li> <li>A single zone can be easily added/updated/deleted at runtime into its live database</li> <li>It’s a known <a href="https://www.knot-dns.cz/benchmark/" target="_blank" rel="noopener">high performance server out of the box</a></li> </ul><p>In theory then what we want isn’t actually that hard:</p><ul> <li>Whenever a domain is added/updated/deleted, we create/replace/remove a zone file for that domain</li> <li>We tell the DNS server to add/update/drop the corresponding zone</li> </ul><p>The devil turned out to be in many small details and mis-adventures.</p><h2 id="lets-not-unknot-the-knot" tabindex="-1">Lets not unknot the knot</h2><p>First things first. Although Knot as a DNS server works great, we’ve found its name to be a bit annoying. It’s a play on the fact that the most common DNS server on the internet is <a href="https://www.isc.org/bind/" target="_blank" rel="noopener">bind</a> so… knot. Unfortunately, you’ll find yourself at some point saying something <a href="https://www.youtube.com/watch?v=8H1u-zh9dmU#t=0m53" target="_blank" rel="noopener">Bernard Woolley-esque</a> like “it was not obvious that it did not work because knot was not bound to the knot ips”. Which is easier to understand when you read it compared to when you say it. We keep trying to keep discussions sane by explicitly saying “ka-not” whenever we refer to the server/software.</p><h2 id="detecting-zone-changes" tabindex="-1">Detecting zone changes</h2><p>There were two main ways we could do this:</p><ol> <li>Catch everywhere in application code we insert/update/delete records for the <code>Domains</code>, <code>CustomDNS</code> or <code>DKIMRecords</code> tables</li> <li>Use DB triggers to do the same thing</li> </ol><p>We ended up going with (2) because it felt like the right choice for rock solid data reliability.</p><p>We update the DB for domains related changes in a number of places, from code, to scripts, to cron jobs. Each of those might need DNS zones to be updated and if we miss something unexpected now or in the future, it might create subtle bugs such as “domain DNS didn’t get updated” or worse, “domain SOA serial number didn’t get bumped and so the Knot replica server didn’t get the updated domain data, so we’re serving different DNS from two different nameservers”.</p><p>These are the sorts of problems that could cause really hard to debug customer issues, that then randomly disappear when the domain gets touched in some other way and everything gets updated correctly again, making them extremely hard to track down and debug.</p><p>DNS is already hard enough for many customers to understand. We wanted to be 100% sure that our DNS servers are rock solid and the data they are generating is completely consistent with what users see in our UI.</p><p>Now actually getting the triggers working turned out to have a number of issues and unexpected edge cases, but also ended up with some nice results.</p><ol> <li>We moved the logic that bumps an SOA serial number into the trigger. Originally this was done on the “active primary” server, and we had to make sure this happened before the non-active failover primary rebuilt the zone. By doing this in the trigger, we ensure that the serial can never be out of date after a change.</li> <li>We ended up with a nice design to track changed domains.</li> </ol><p>We have a <code>KnotDomainsChanged</code> table that looks like:</p><pre><code>+--------------+--------------+------+-----+---------+-------+ 258 | Field | Type | Null | Key | Default | Extra | 259 +--------------+--------------+------+-----+---------+-------+ 260 | Server | varchar(255) | NO | PRI | NULL | | 261 | DomainId | int | NO | PRI | NULL | | 262 | Domain | varchar(255) | YES | | NULL | | 263 | NeedsRebuild | tinyint | YES | | 1 | | 264 | ErrorCount | int | YES | | 0 | | 265 +--------------+--------------+------+-----+---------+-------+ 266 </code></pre><p>There is another table <code>KnotServers</code> that has a list of all currently running Knot servers. Whenever a domain is added/updated/deleted, a trigger executes this query:</p><pre><code> INSERT INTO KnotDomainsChanged (Server, DomainId, Domain) 267 SELECT Server, BumpDomainId, BumpDomain 268 FROM KnotServers 269 ON DUPLICATE KEY UPDATE 270 NeedsRebuild = 1, 271 ErrorCount = 0; 272 </code></pre><p>With this, the <code>KnotDomainsChanged</code> table effectively maintains a set of changed domains for each primary Knot server to pick up and build. By using <code>(Server, DomainId)</code> as the primary key, we ensure this table can’t grow without bound even if a primary Knot server is down for a while.</p><p>The <code>NeedsRebuild</code> flag allows us to correctly rebuild a zone without a race condition. The process for keeping zones up-to-date is to effectively run the following in an infinite loop.</p><ul> <li>Fetch and iterate over all <code>KnotDomainsChanged</code> records for this Server <ul> <li>Set <code>NeedsRebuild = 0</code> for this <code>(Server, DomainId)</code> record</li> <li>If the domain exists in the <code>Domains</code> table, add/rebuild the zone into Knot</li> <li>If the domain does not exist in the <code>Domains</code> table, purge the zone from Knot</li> <li>Delete the <code>(Server, DomainId)</code> record from <code>KnotDomainsChanged</code> iff <code>NeedsRebuild = 0</code></li> </ul> </li> </ul><p>So if the domain changes while a rebuild is in progress, the trigger will set <code>NeedsRebuild = 1</code>, which means it won’t be deleted from the table after the zone build finishes, which means it’ll be picked up again the next time the sync runs again in a few seconds. This avoids the race if the domain is changed while it is in the process of being rebuilt.</p><h2 id="using-standard-dns-axfr-ixfr-for-replication" tabindex="-1">Using standard DNS AXFR/IXFR for replication</h2><p>Our initial plan was to use a more “standard” DNS setup. We would have an internal hidden primary server and a number of secondary servers that pull from the primary via <a href="https://datatracker.ietf.org/doc/html/rfc5936" target="_blank" rel="noopener">AXFR</a>/IXFR. Only the secondary servers would handle DNS queries from the world.</p><p>This turned out to have a number of annoying edge cases:</p><ol> <li>It required two completely separate Knot server setups with completely separate and quite different configurations. Although this is a more “standard” DNS management approach, it reduced some of benefit of moving to a single DNS server.</li> <li>Using AXFR for updates caused unexpected problems. AXFR connections aren’t reused, so every domain transferred required a separate TCP connection. When a lot of domains needed to be updated at once, we <a href="https://vincent.bernat.ch/en/blog/2014-tcp-time-wait-state-linux" target="_blank" rel="noopener">ran out of TCP socket tuples because of sockets in TIMEWAIT state</a>. This caused AXFRs to the secondaries to start failing. We fixed this by enabling the kernel <code>tcp_tw_reuse</code> tunable, but it felt… hacky.</li> <li>At Fastmail, whenever we have a singleton service (e.g. Knot primary), we want to make sure that we can take the machine it’s running on down safely. To allow that we have the service run on at least two separate servers and use a failover IP to bind to the current up/active server.</li> </ol><p>Unfortunately this combined with the way Knot does catalog zones completely broke AXFR/IXFR replication. What is a <a href="https://datatracker.ietf.org/doc/rfc9432/" target="_blank" rel="noopener">catalog zone</a>? It’s a standard way to allow primary and secondary DNS servers to keep the complete list of zones actually managed in sync.</p><pre><code> The content of a DNS zone is synchronized among its primary and 273 secondary nameservers using AXFR and IXFR. However, the list of 274 zones served by the primary (called a catalog in [RFC1035]) is not 275 automatically synchronized with the secondaries. To add or remove a 276 zone, the administrator of a DNS nameserver farm not only has to add 277 or remove the zone from the primary, they must also add/remove 278 configuration for the zone from all secondaries. This can be both 279 inconvenient and error-prone; in addition, the steps required are 280 dependent on the nameserver implementation. 281 </code></pre><p>It’s a classic example of taking a system that already has a way of storing data (DNS records) and replicating that data (AXFR/IXFR), and reusing those mechanisms to sync something else. In this case, a specially configured catalog zone that itself contains a list of all other member zones managed by the server in a standard defined format.</p><p>The basic format of the catalog zone is that each member zone managed by the server exists in the catalog zone as the RDATA of a PTR record. Then because each PTR record needs to be a unique domain name, a unique identifier is used as a sub-domain of the catalog zone itself. e.g. <code>unique-N.catalog.zone. PTR member-domain.org.</code></p><p>The problem here is that the unique identifiers are not generated in a consistent way if you have multiple different primary servers!</p><pre><code>k1: 8815a670d8fa9032.zones.catalog.dns.internal. 0 PTR uberengineer.com. 282 k2: 09b6b01fb75e81e6.zones.catalog.dns.internal. 0 PTR uberengineer.com. 283 </code></pre><p>During testing k1 was the entry on our first server (e.g. active primary), the k2 on the second server (e.g. backup primary).</p><p>So when we did a failover from the current active primary server to promote the backup primary to the active primary, the catalog zone on the new active primary is completely out of sync with the catalog zone on all the downstream secondary servers. This caused a massive amount of resyncing and general Knot confusion. There didn’t seem to be an easy solution to this problem without ultimately having some singleton source of truth server, which is what we wanted to avoid.</p><h2 id="switching-to-a-single-server-type" tabindex="-1">Switching to a single server type</h2><p>After this, we decided to dump the whole Knot primary/secondary system and AXFR/IXFR replication, and instead have just a single type of Knot server. This ended up having a number of advantages.</p><ol> <li>Only one type of Knot server and one type of Knot server configuration. Less configurations to understand.</li> <li>No catalog zone. Less concepts to learn about.</li> <li>No AXFR/IXFR. AXFR/IXFR is harder to reason about and less visible to operators.</li> <li>No worry about zone serial numbers going backwards if you add -&gt; update -&gt; delete -&gt; re-add a domain which can cause AXFR/IXFR replication weirdness.</li> </ol><p>Doing this definitely felt like a better solution. Additionally it already all “just worked” because we had already built the trigger system that could keep an arbitrary number of primary servers (originally for failover) up-to-date. Now it was just keeping all our Knot instances up to date.</p><h2 id="testing-the-new-system" tabindex="-1">Testing the new system</h2><p>Since DNS is so critical and we have 100,000’s of domains, we wanted to test as carefully as possible that the new Knot system would generate the same results as the existing PowerDNS system.</p><p>The <code>Net::Pcap</code>, <code>Net::Frame</code> and <code>Net::DNS</code> modules in perl made this straight forward. We were able to write a script that captured packets with libpcap, unpacked them into DNS queries and responses, and then replayed them against the new Knot servers to compare the results. This allowed us to see with real world query data that we would get back the same responses.</p><pre><code>my $resolver = Net::DNS::Resolver-&gt;new(nameservers =&gt; [ '...existing DNS ip...' ]); 284 ... libpcap setup ... 285 my $link_class; 286 if ($linktype == Net::Pcap::DLT_EN10MB) { 287 $link_class = 'Net::Frame::Layer::ETH'; 288 } elsif ($linktype == Net::Pcap::DLT_LINUX_SLL) { 289 $link_class = 'Net::Frame::Layer::SLL'; 290 } else { 291 die &quot;unknown link layer: $linktype\n&quot;; 292 } 293 ... run libpcap loop ... 294 sub process_packet ($user_data, $header, $packet) { 295 my $p_link = $link_class-&gt;new(raw =&gt; $packet); 296 $p_link-&gt;unpack; 297 my $p_ip4 = Net::Frame::Layer::IPv4-&gt;new(raw =&gt; $p_link-&gt;payload); 298 $p_ip4-&gt;unpack; 299 my $p_udp = Net::Frame::Layer::UDP-&gt;new(raw =&gt; $p_ip4-&gt;payload); 300 $p_udp-&gt;unpack; 301 my $p_dns = Net::DNS::Packet-&gt;decode( \$p_udp-&gt;payload ); 302 303 if (my @a = $p_dns-&gt;answer) { 304 my @q = $p_dns-&gt;question; 305 my $q = $q[0]; 306 307 my $dns_q = Net::DNS::Packet-&gt;new(); 308 $dns_q-&gt;push(question =&gt; $q); 309 my $kres = $resolver-&gt;send($dns_q); 310 my @ka = $kres-&gt;answer; 311 312 ... compare @a (pdns answer) vs @ka (knot answer) modulo some known differences ... 313 </code></pre><p>Mostly it showed that everything was working as expected, though there were a few interesting edge cases that ended up needing to be dealt with.</p><h2 id="non-terminal-nodes-and-wildcards" tabindex="-1">Non-terminal nodes and wildcards</h2><p>The biggest subtle difference we discovered was around non-terminal nodes and wildcards.</p><p>This is subtly documented in the tinydns <a href="http://cr.yp.to/djbdns/axfr-get.html" target="_blank" rel="noopener">axfr-get</a> program:</p><blockquote> <p>axfr-get does not precisely simulate BIND’s handling of <code>*.dom</code>. Under BIND, records for <code>*.dom</code> do not apply to <code>y.dom</code> or <code>anything.y.dom</code> if there is a normal record for <code>x.y.dom</code>. With axfr-get and tinydns, the records apply to <code>y.dom</code> and <code>anything.y.dom</code> except <code>x.y.dom</code>.</p> </blockquote><p>Knot DNS follows the traditional BIND and <a href="https://datatracker.ietf.org/doc/html/rfc1034#section-4.3.3" target="_blank" rel="noopener">RFC 1034</a> intepretation of wildcard records. Our PowerDNS backend was built to follow the TinyDNS model because that’s what we were migrating from at the time. This can cause subtle differences in DNS results if you have any wildcard domains.</p><p>An example. If you have a domain <code>example.com</code> setup at Fastmail with our standard DNS configuration, then we add default A records for <code>*.example.com</code>. If you add a DNS record for the subdomain <code>*.foo</code> that creates a record for <code>*.foo.example.com</code>. However that implied existence of <code>foo.example.com</code> means that the <code>*.example.com</code> record no longer exists for <code>foo.example.com</code>. At least, that’s true in the Knot/Bind DNS implementation, but not true in our PowerDNS backend implementation, so this will get you subtly different results.</p><p>In quite a few cases this won’t be a problem. We see queries like <code>_domainkey.example.com A</code>, which under PowerDNS return an IP, but won’t under knot, but that’s fine. Nothing should really be using that IP anyway, it’s probably just some gateway device somewhere doing DNS querying of passively seen domains in email headers or the like. But in some cases it might be.</p><p>We initially thought we could fix this automatically. For everyone that has a <code>foo.bar.example.com</code> subdomain without a <code>bar.example.com</code> subdomain, if there’s any wildcard <code>*.example.com</code> records, we make copies of them at <code>bar.example.com</code>. This should make everything “just work”.</p><p>The problem is, there’s actually a deeper problem that affects basically every domain and every single subdomain in a subtle way. As noted, by default we publish <code>*.example.com</code> records. However these effectively work for all sub-sub domains as well. For example I have a domain <code>uberengineer.com</code> and the standard <code>*.uberengineer.com</code> wildcard means that all sub-domains, sub-sub-domains, etc resolve:</p><pre><code># dig +short that.uberengineer.com 314 103.168.172.37 315 103.168.172.52 316 # dig +short this.that.uberengineer.com 317 103.168.172.37 318 103.168.172.52 319 </code></pre><p>Now I have a TXT record at <code>test.uberengineer.com</code>, so that hides the A records</p><pre><code># dig +short test.uberengineer.com 320 # 321 </code></pre><p>But in our PowerDNS implementation, that doesn’t hide sub-domains of <code>test.uberengineer.com</code> from the original wildcard.</p><pre><code># dig +short foo.test.uberengineer.com 322 103.168.172.37 323 103.168.172.52 324 </code></pre><p>But with Knot, it does:</p><pre><code># dig +short foo.test.uberengineer.com @knottest.internal 325 # 326 </code></pre><p>So basically to just fix users automatically, for every single subdomain X they have configured, if there’s a wildcard at a lower level, we’d have to explicitly create a copy of all the lower level wildcard records at <code>*.X</code> as well. This was going to be just way too much magic, especially for a rare edge case that it’s possible no one was even relying on anyway!</p><p>In the end, we analysed the DNS queries we were seeing and also all the email deliveries we saw. We can looked at the email logs on our MX servers for any RCPT TO address, and then checked if the domain that was delivered to is hosted by us, and compared if it would resolve under Knot as well. Combining the data convinced us that no one was going to be actively affected by this change.</p><h2 id="conclusion" tabindex="-1">Conclusion</h2><p>This has now been running for over a year in production and has been working extremely well. We were able to remove two existing systems and replace them with a single consistent Knot DNS based system for all our Fastmail domains and 100,000’s of user domains. The triggers that track zones that need rebuilding work reliably and consistently. The new system performs enormously better than the existing PowerDNS pipe backend based system. We can easily scale it to add additional servers if needed.</p><p>This all fits with a mantra we’ve been working with recently, “fewer better ways”. We’ve been running an email service for over 25 years and it’s easy to accumulate a plethora of different services and systems with varying levels of polish and performance. Revisiting what you’re doing and running to try and consolidate to a smaller number of systems working in better ways can reduce long term debt. This makes systems more reliable, easier to manage, and also allows new staff to understand them more quickly as well.</p></content> 327 </entry><entry> 328 <title>Dec 12: Following the Sun</title> 329 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/following-the-sun/' /> 330 <id>https://www.fastmail.com/blog/following-the-sun/</id> 331 <updated>2024-12-12T00:00:01Z</updated><author> 332 <name>Andria DeFulio</name> 333 </author><author> 334 <name>Luke Erlacher</name> 335 </author><content xml:lang='en' type='html'><p>This is the twelfth post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/meet-the-team-marc/">Dec 11: Meet the team—Marc</a>. The next post is <a href="/blog/moving-fastmail-dns-to-knot/">Dec 13: It’s knot DNS. There’s no way it’s DNS. It is DNS!</a>.</p><p>As Fastmail provides a global service for end-users, we operate 24/7 and respond to issues around the clock. Given that humans do not operate around the clock, this requires some thoughtful process to effectively maintain our service.</p><p>To do this, our support team is split across three timezones, and our on-call engineers are split across two timezones.</p><p>As a term of art, this is often referred to as “Follow the Sun support”: Wherever the sun is shining right now, agents and support engineers are working and ready there, resulting in 24/7 coverage.</p><h2 id="support" tabindex="-1">Support</h2><p>Our support staff is located in both of our offices in Melbourne and Philadelphia, and remotely in India. Having staff located around the world works really well for us and for our customers, who are also spread across the globe!</p><p>Having staff spread across multiple time zones allows us to provide support to you, our customer, when you need it. We aim to respond to routine questions within a few hours. We usually do much better than that, and respond within an hour!</p><p>If we were all located in a single office, not only would our customers potentially be left waiting for a full day to get a response to a ticket, but our support team would start each morning with a daunting backlog of tickets. It’d be much trickier to surface customers’ most urgent tickets, and the natural response to a backlog might be to rush through tickets. Instead, we have opted out of that unnecessary stress and can give each user the time it takes to fully research their issue and then send them a thorough, thoughtful response.</p><p>Having a team across multiple locations does present challenges, too! Particularly with collaboration. To mitigate that, the support team has multiple brief huddles, with each location having some hours of overlap within their schedule. At the end of our day, we hand the baton off to those who are just starting their day, letting them know if there is anything impacting multiple customers, like <a href="/blog/moving-house-new-datacentre/">moving to a new data center</a> or <a href="/blog/sunsetting-pobox/">migrating Pobox users</a>. We also use these huddles to help each other with particularly challenging tickets, to make sure the team is aware of work on the horizon, and to share newly acquired knowledge across the team.</p><p>Outside of our daily huddles, we communicate with everyone on the Fastmail team synchronously in Slack and asynchronously via our sister product, <a href="https://www.topicbox.com/" target="_blank" rel="noopener">Topicbox</a>.</p><h2 id="on-call-engineers" tabindex="-1">On-Call Engineers</h2><p>At Fastmail, we own and operate most of our stack, from networking gear and servers in the datacenter space we use, server OS setup, internal services, backups, and redundancy, through to user-facing services and mail flow.</p><p>To maintain the availability of all these 24/7, our platform team operates an on-call process split between our US and Australian teams.</p><p>With the Australian team working to Australia East timezone, and the US team working to Eastern Standard timezone, we have set this up for the Australian team to be on call from 8AM to 10PM their time, and the US team from 6AM to 4PM their time.</p><p>So the times are a little less friendly for the US team, but in return, they have less time to cover overall.</p><p>As well as being the primary on-call engineers, the platform team also owns our observability stack and is empowered to make changes to all parts of our stack to improve monitoring and alerting. Nothing is worse than being responsible for something you can’t fix, so we make sure to avoid that!</p><p>Another practice we follow to reduce on-call load is to help the support team troubleshoot and fix issues without needing to escalate to engineers. Our support team can search logs and run admin and fixup tools to resolve problems themselves.</p><p>One of Fastmail’s operational maxims is “the spice must flow”. So, the first priority during incidents is to restore service—the mail has to flow, and users need to be able to access their mailboxes and use all the other features we provide. Our on-call engineers are experts in most, but not all, systems and processes at Fastmail, so sometimes this involves paging experts or technical owners to help troubleshoot.</p><p>Once service is restored, everything else has less urgency, and incident reports and follow-ups are coordinated via - what else - emails on our internal mailing lists. We make sure that we understand what happened, fix bugs, and make the system more robust for the future.</p></content> 336 </entry><entry> 337 <title>Dec 11: Meet the team—Marc</title> 338 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/meet-the-team-marc/' /> 339 <id>https://www.fastmail.com/blog/meet-the-team-marc/</id> 340 <updated>2024-12-11T00:00:01Z</updated><author> 341 <name>The Fastmail Team</name> 342 </author><content xml:lang='en' type='html'><p>This is the eleventh post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/sunsetting-pobox/">Dec 10: Sunsetting Pobox</a>. The next post is <a href="/blog/following-the-sun/">Dec 12: Following the Sun</a>.</p><p>Meet Marc, our Head of the new Trust, Abuse, and Deliverability team — and also source of bad puns. (Oy, good puns!)</p><p><strong>Name:</strong> Marc Bradshaw</p><p><strong>Role:</strong> Head of Trust, Abuse, and Deliverability</p><p><strong>What do you work on?</strong></p><p>The Trust, Abuse, and Deliverability (TAD) team is responsible for mail flow, and anti-abuse. It is our task to make sure the good mail gets delivered and the bad mail does not. This is a fairly new team here at Fastmail, with the responsibility having previously been split between the backend and platform teams.</p><p><strong>How long working at Fastmail, how did you get involved?</strong></p><p>Too bloody long! 10 years now. I started in 2014. Rob Norris called me and asked if I was still looking for a job. I said “no, but keep talking”. I was another refugee from Monash University - we had worked there together. I had started somewhere else, but that was not a good fit.</p><p><strong>What’s a project you have worked on that you’re particularly proud of?</strong></p><p>The answer to this is quite often “the next one”, we are rewriting our inbound mail delivery software, and have lots of great ideas which should allow some great new features in the future.</p><p><strong>What are your favourite Fastmail features?</strong></p><p>I love the memos feature. Being able to put a note on an email so you can remember something about it. Also, labels! Labels make organising email easy.</p><p><strong>Other than Fastmail, what’s your favourite or most used piece of technology?</strong></p><p>I have a supernote nomad, which is an e-ink tablet which I use to take and organise notes, and as an e-reader. Being e-ink allows me to write notes in a more natural way, but still be able to organise them digitally, all without the distractions I would get from a more fully featured tablet.</p><p><a href="https://nothingbutstatic.dev/tech/2024-07-13-the-digital-analog-2/" target="_blank" rel="noopener">I blogged about it</a> back in April.</p><p><strong>What are you listening to these days?</strong></p><p>My musical taste is quite varied, my recently played list on the music app ranges from 80s synth pop, to Madonna, Linkin Park, SOPHIE, Veruca Salt, and Chappell Roan</p><p><strong>What are you watching these days?</strong></p><p>I have just finished watching Fallout, I can’t believe we need to wait until 2026 for the next season.</p><p><strong>What do you like to do outside of work?</strong></p><p>I like to ride my bike and try to keep fit, with varying levels of success.</p><p><strong>What’s your favourite animal?</strong></p><p>Hmm. Just got a new puppy. He’s pretty cute, when he’s not being naughty and getting stuck under the house!</p><p><strong>Any Fastmail staff you want to brag on?</strong></p><p>I mean, they’re all pretty good. Andrew has moved from the Platform team into Development and is quickly getting up to speed on all things backend and mail flow.</p><p><strong>What do you like best about working at Fastmail?</strong></p><p>I love that we are a small company making a big difference in our field.</p></content> 343 </entry><entry> 344 <title>Dec 10: Sunsetting Pobox</title> 345 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/sunsetting-pobox/' /> 346 <id>https://www.fastmail.com/blog/sunsetting-pobox/</id> 347 <updated>2024-12-10T00:00:01Z</updated><author> 348 <name>Ricardo Signes</name> 349 </author><content xml:lang='en' type='html'><p>This is the tenth post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/building-a-blog/">Dec 9: Building a blog</a>. The next post is <a href="/blog/meet-the-team-marc/">Dec 11: Meet the team—Marc</a>.</p><p>On November 12th of this year, our service Pobox was finally merged into Fastmail. This had been inevitable since <a href="/blog/exciting-news-about-pobox-and-fastmail/">Fastmail’s acquisition of Pobox</a> in 2015, but we’d put it off over and over. In the end, it took a lot more time than we expected, but it paid off. The cutover went well… but it was still just a little bittersweet, ending a long era. Pobox was around for thirty years. In Internet years, that’s an eternity. Seeing something so long-lasting cease to exist as itself is a bit, well, <em>confronting</em>.</p><p>It hit home for me, personally: I’d been working on Pobox for nearly 20 years, and the shutdown of Pobox meant turning off hundreds of thousands of lines of code I wrote or maintained. A professional programmer does well to avoid identifying too much with their work product, but on some level it’s hard to avoid. Still, it was better to do it myself. I had a good knowledge of both systems and felt like I was best placed to make things as seamless as possible.</p><p>So, how’d it all go? I’ll walk you through it.</p><h2 id="what-even-was-pobox" tabindex="-1">What even was Pobox?</h2><p>Pobox was, at its heart, an email forwarding service. It was started by Meng Wong and Helen Horstmann-Allen in 1995, while they were both still in college. College students in the 90s had email addresses, but they knew they wouldn’t last. Eventually you’d leave school and you’d need to tell everybody your new address. The same thing went for people who had email at work. Pobox was the answer to this problem: you’d get an address with Pobox and set it up to forward your mail to whatever your current “real” address was. Unlike Fastmail, Pobox didn’t include mail storage by default. You could get it, but it was an add-on.</p><p>Broadly speaking, Pobox and Fastmail were very similar. Users would have one or more email addresses, and those addresses would deliver to an inbox or to another address. Spam got filtered out, and users could write their own mail filters. One “group” could have a bunch of users on it. Users could send mail through relay servers.</p><h2 id="so-what-was-the-problem" tabindex="-1">So what was the problem?</h2><p>Yes, broadly speaking, Pobox and Fastmail <em>were</em> very similar. In the fine details, though, everything was just kind of different in tedious little ways. Email aliases were per-group in Fastmail, per-user in Pobox. Spam filtering exceptions worked differently. The rules on sharing domains were different. In short, no feature <em>really</em> worked the same in both services, and every exception had an exception.</p><p>Obviously, maintaining many of the same features (but differently) in two places took a lot more time and effort than just maintaining one version. The lion’s share of development always went to Fastmail while the number of staff with expertise in Pobox dwindled. We kept training new support staff on Pobox, but new programmers almost never touched it. Pobox had emergency support staff, but no full time developers, and it showed, with nearly no changes of any kind after 2015. That didn’t just mean they weren’t getting cool new features. It meant they weren’t getting as many deliverability improvements or updates to service monitoring.</p><p>In order to keep Pobox users happy, online, and connected, we were going to have to turn them into Fastmail users. We wanted to make this as seamless as possible, which was going to be real work. Remember: everything was different enough to make “just import them to Fastmail” not work. Still, the goal was that cutover would happen without significant downtime, and without users needing to reconfigure their mail clients, change passwords, or (worst of all) fiddle with their DNS.</p><p>This was going to be a complicated piece of work, and risky, and was going to require significant knowledge of both Pobox and Fastmail internals. Also, it just wasn’t going to be exciting work day to day. The final payoff would be great, but otherwise it would involve writing a lot of fiddly and risky code that was just going to get deleted later. So we put it off a long time, even though we’d known since 2015 that we’d have to do it eventually. In 2023, we finally started in earnest.</p><h2 id="the-upgrade-process" tabindex="-1">The Upgrade Process</h2><p>Back when Fastmail acquired Pobox, the first thing we changed was the Pobox mail storage. Both Fastmail and Pobox used <a href="https://github.com/cyrusimap/cyrus-imapd" target="_blank" rel="noopener">Cyrus IMAP</a>, but Fastmail ran it at larger scale and with better tooling. Very early on, we created a secret Fastmail user for each Pobox Mailstore account and moved mail out of Pobox and into Fastmail. This also let us replace Pobox’s webmail with Fastmail’s best-in-class UI.</p><p>It didn’t replace almost anything else, though: IMAP and SMTP connections went through Pobox. Pobox handled the mail forwarding. Settings and billing were all in Pobox. That meant that those “secret Fastmail users” had all kinds of settings that nothing used. The Pobox Upgrade Project (known internally as PUP) used these as the starting point: they would be configured to act almost exactly like their Pobox owners, and then they’d take over.</p><p>First, we created these Fastmail users for all the non-Mailstore Pobox accounts. This was pretty painless. Next, we merged these users into customers to match the billing setup in Pobox. That’s where the pain began. Fastmail supported merging multiple users into one customer (to get just one bill)… but not when both of those users are free and have no expiration date. We had to get into the billing system and add ways to override the safety mechanisms keeping us from breaking the rules that we so desperately needed to break. I won’t provide a litany of every such complication, but: there were plenty. Each ported feature involved one or two big mismatches, one or two edge cases, and one or two places where everything would work just fine, if you just bypassed all the business logic briefly.</p><p>For months, the process moved forward feature by feature. At Pobox, the “migration planner” kept learning how to describe more and more of each Pobox account’s configuration in terms of Fastmail. The planner wrote out new plans constantly, and those were shipped from Pobox to Fastmail. On the Fastmail side, the “migration executor” would take those plans and reconfigure things on the matching Fastmail users. Every week, the Fastmail users were configured more and more like their counterpart Pobox users. The big question was: when could we cut over?</p><h2 id="cutover" tabindex="-1">Cutover</h2><p>When making a big change, we generally like a staged rollout. First, we apply it to some of our internal test customers. Then we apply it to ourselves. Then one percent of users, then five… you get the idea. We’re looking to reduce the damage that a bug can cause. The Pobox cutover was exactly the kind of big (huge!) change where we’d like roll things out in stages.</p><p>Unfortunately, it wasn’t going to be that simple. Email is routed by domain. You can’t really send one user’s mail through one system and another user’s mail through another mail system. Or, you can, but it requires building a <em>third</em> system that sits in front of both, which is just another thing that can go wrong!</p><p>We decided that all users would get cut over at once. To mitigate that big change, we further decided that we’d cut over different components over time, as much as possibly invisibly to the user, like this:</p><ul> <li>First, anybody connecting to IMAP or POP was connecting directly to Fastmail, not Pobox.</li> <li>Later, DNS for our domains (and user domains) was moved to Fastmail.</li> <li>Later still, anybody sending mail through Pobox was actually connecting directly to Fastmail, not Pobox.</li> </ul><p>Each of these cutovers was a big change that we’d been able to test well in advance, and each one turned up one or two weird edge cases that took a little time to fix. Doing them one at a time let us focus and get things right. As far as I know, nobody reported noticing that any service had moved to Fastmail.</p><p>Finally, the big cutover day was set for November 12th. That’s when we’d have all mail destined for Pobox customers start hitting Fastmail servers, and when we’d make <a href="https://pobox.com/" target="_blank" rel="noopener">pobox.com</a> start taking people to Fastmail. We had a big checklist, and the plan was that we’d start our day at 8:00, run a series of programs, watch some logs, and be done around lunchtime. That was the prediction if everything went <em>perfectly</em>. In reality, we weren’t done until about three in the afternoon, and we did spend a decent amount of time responding to unexpected cases. Mostly, though, it went smoothly and we were done by the end of the day.</p><p>The goal had been to avoid a painful reconfiguration or transition process for users, and I think we did pretty well on that front! I’m proud of that. Pobox is how I got involved in working on email, and it’s been one of the major things I’ve worked on for twenty years. Working on Pobox made me a better programmer and helped me understand what it means to be a good internet citizen and a good steward of customers’ data. Astoundingly, most Pobox customers have been with Pobox <em>at least</em> as long as I have. I’m pleased we could bring them smoothly into Fastmail, where we can keep trying to deliver the best email service around.</p><p>Everybody at Fastmail contributed to the success of the project, but I think I should especially shout out two teams. Team PUP (meaning Mark Jason Dominus and Matthew Horsfall) worked on the project full time, dealing with weird legacy systems, obnoxious edge cases, and plenty of complexity arising from wiring together two incompatible systems. It was a bit of a slog, and they did great. Also, our support team (Team SUP!) were invaluable in this project. They wrote the <a href="https://www.fastmail.help/hc/en-us/articles/9822848635919-Guide-to-Fastmail-for-Pobox-users" target="_blank" rel="noopener">Guide to Fastmail for Pobox users</a>, they provided feedback on what decisions would help make things easiest on users (and on them!), and they took care of our users whenever things weren’t <em>quite</em> as seamless as we’d hoped. One of the most common concerns I heard from Pobox customers was “the support won’t be as good.” The great news is that it’s just one support team that’s been supporting both products, so the same great support is going to be there for you.</p><p>Thanks for a few good decades. I’m looking forward to what comes next!</p></content> 350 </entry><entry> 351 <title>Dec 9: Building a blog</title> 352 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/building-a-blog/' /> 353 <id>https://www.fastmail.com/blog/building-a-blog/</id> 354 <updated>2024-12-09T02:00:01Z</updated><author> 355 <name>Callum Skeet</name> 356 </author><content xml:lang='en' type='html'><p>This is the ninth post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/principles/">Dec 8: Guiding principles</a>. The next post is <a href="/blog/sunsetting-pobox/">Dec 10: Sunsetting Pobox</a>.</p><p>Today we’ll look at how we went about creating our new blog and marketing site, why we chose the tools we did, and how we resolved issues we bumped into along the way.</p><h2 id="moving-away-from-word-press" tabindex="-1">Moving away from WordPress</h2><blockquote> <p>N.B. Eleventy now <a href="https://github.com/11ty/eleventy-import" target="_blank" rel="noopener">provides a tool</a> to assist with migrating content from WordPress</p> </blockquote><p>Our last site was powered by WordPress and, despite its age, it also powers a <a href="https://w3techs.com/technologies/details/cm-wordpress" target="_blank" rel="noopener">large proportion</a> of the websites out there. Just because something’s popular doesn’t mean it’s the right tool for us. The added costs of running a WordPress instance aren’t insignificant, and making large, sweeping changes to any aspect of the site required a level of WordPress knowledge that few people wanted to learn (or admit they knew). So, what we wanted was something simple, that anyone can contribute to, and considerably lowered our running costs.</p><h2 id="meet-eleventy" tabindex="-1">Meet Eleventy</h2><p><a href="https://www.11ty.dev/" target="_blank" rel="noopener">Eleventy (11ty)</a> is a Static Site Generator (SSG). It’s also a tool that fit our needs perfectly. It was written with Node (1 fewer language to know), allowed us to write content in markdown, and spat out HTML which can be hosted anywhere. It’s also really fast. Adding extra functionality to it is also dead simple.</p><p>For example, we can easily add a new font family like so:</p><pre class="language-jsx"><code class="language-jsx"><span class="comment token">// -----------------</span> 357 <span class="comment token">// source/_data/webfonts.js</span> 358 <span class="comment token">// -----------------</span> 359 360 <span class="keyword token">const</span> config <span class="operator token">=</span> <span class="punctuation token">{</span> 361 <span class="comment token">// ...</span> 362 <span class="literal-property property token">roca</span><span class="operator token">:</span> <span class="punctuation token">[</span> 363 <span class="punctuation token">{</span> <span class="literal-property property token">weight</span><span class="operator token">:</span> <span class="constant token">THIN</span><span class="punctuation token">,</span> <span class="literal-property property token">name</span><span class="operator token">:</span> <span class="string token">'th'</span> <span class="punctuation token">}</span><span class="punctuation token">,</span> 364 <span class="punctuation token">{</span> <span class="literal-property property token">weight</span><span class="operator token">:</span> <span class="constant token">LIGHT</span><span class="punctuation token">,</span> <span class="literal-property property token">name</span><span class="operator token">:</span> <span class="string token">'lt'</span> <span class="punctuation token">}</span><span class="punctuation token">,</span> 365 <span class="punctuation token">{</span> <span class="literal-property property token">weight</span><span class="operator token">:</span> <span class="constant token">REGULAR</span><span class="punctuation token">,</span> <span class="literal-property property token">name</span><span class="operator token">:</span> <span class="string token">'rg'</span> <span class="punctuation token">}</span><span class="punctuation token">,</span> 366 <span class="punctuation token">]</span><span class="punctuation token">,</span> 367 <span class="punctuation token">}</span><span class="punctuation token">;</span> 368 369 <span class="keyword token">function</span> <span class="function token">getRocaVariant</span><span class="punctuation token">(</span><span class="parameter token"><span class="punctuation token">{</span> weight<span class="punctuation token">,</span> name <span class="punctuation token">}</span></span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 370 <span class="keyword token">return</span> <span class="punctuation token">{</span> 371 <span class="property string-property token">'@font-face'</span><span class="operator token">:</span> <span class="punctuation token">{</span> 372 <span class="property string-property token">'font-family'</span><span class="operator token">:</span> <span class="string token">'roca'</span><span class="punctuation token">,</span> 373 <span class="property string-property token">'src'</span><span class="operator token">:</span> <span class="template-string token"><span class="string template-punctuation token">`</span><span class="string token">url(/assets/fonts/roca/rocaone-</span><span class="interpolation token"><span class="interpolation-punctuation punctuation token">${</span>name<span class="interpolation-punctuation punctuation token">}</span></span><span class="string token">-webfont.woff2) format('woff2')</span><span class="string template-punctuation token">`</span></span><span class="punctuation token">,</span> 374 <span class="property string-property token">'font-weight'</span><span class="operator token">:</span> weight<span class="punctuation token">,</span> 375 <span class="property string-property token">'font-style'</span><span class="operator token">:</span> <span class="string token">'normal'</span><span class="punctuation token">,</span> 376 <span class="property string-property token">'font-display'</span><span class="operator token">:</span> <span class="string token">'swap'</span><span class="punctuation token">,</span> 377 <span class="punctuation token">}</span><span class="punctuation token">,</span> 378 <span class="punctuation token">}</span><span class="punctuation token">;</span> 379 <span class="punctuation token">}</span> 380 381 <span class="keyword token">export</span> <span class="keyword token">default</span> <span class="punctuation token">{</span> <span class="literal-property property token">roca</span><span class="operator token">:</span> config<span class="punctuation token">.</span>roca<span class="punctuation token">.</span><span class="function token">map</span><span class="punctuation token">(</span>getRocaVariant<span class="punctuation token">)</span> <span class="punctuation token">}</span> 382 383 <span class="comment token">// -----------------------------</span> 384 <span class="comment token">// source/_includes/layouts/base.liquid</span> 385 <span class="comment token">// -----------------------------</span> 386 387 <span class="tag token"><span class="tag token"><span class="punctuation token">&lt;</span>style</span><span class="punctuation token">></span></span><span class="plain-text token"> 388 {%- toCSS from: webfonts.roca -%} 389 </span><span class="tag token"><span class="tag token"><span class="punctuation token">&lt;/</span>style</span><span class="punctuation token">></span></span> 390 391 <span class="comment token">// ------------</span> 392 <span class="comment token">// .eleventy.js</span> 393 <span class="comment token">// ------------</span> 394 395 eleventyConfig<span class="punctuation token">.</span><span class="function token">addPassthroughCopy</span><span class="punctuation token">(</span><span class="string token">'source/assets/fonts'</span><span class="punctuation token">)</span><span class="punctuation token">;</span> 396 eleventyConfig<span class="punctuation token">.</span><span class="function token">addLiquidTag</span><span class="punctuation token">(</span><span class="string token">'toCSS'</span><span class="punctuation token">,</span> <span class="keyword token">function</span> <span class="punctuation token">(</span><span class="comment token">/* liquidEngine */</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 397 <span class="keyword token">return</span> <span class="punctuation token">{</span> 398 <span class="function token">parse</span><span class="punctuation token">(</span><span class="parameter token">tagToken</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 399 <span class="keyword token">this</span><span class="punctuation token">.</span>args <span class="operator token">=</span> <span class="keyword token">new</span> <span class="class-name token">Hash</span><span class="punctuation token">(</span>tagToken<span class="punctuation token">.</span>args<span class="punctuation token">)</span><span class="punctuation token">;</span> 400 <span class="punctuation token">}</span><span class="punctuation token">,</span> 401 <span class="operator token">*</span><span class="function token">render</span><span class="punctuation token">(</span><span class="parameter token">context<span class="punctuation token">,</span> emitter</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 402 <span class="keyword token">const</span> <span class="punctuation token">{</span> from <span class="punctuation token">}</span> <span class="operator token">=</span> <span class="keyword token">yield</span> <span class="keyword token">this</span><span class="punctuation token">.</span>args<span class="punctuation token">.</span><span class="function token">render</span><span class="punctuation token">(</span>context<span class="punctuation token">)</span><span class="punctuation token">;</span> 403 <span class="comment token">// Always a good idea to cache processed CSS/JS where you can.</span> 404 <span class="comment token">// This can speed your builds up significantly.</span> 405 <span class="keyword token">if</span> <span class="punctuation token">(</span>cssCache<span class="punctuation token">.</span><span class="function token">has</span><span class="punctuation token">(</span>from<span class="punctuation token">)</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 406 emitter<span class="punctuation token">.</span><span class="function token">write</span><span class="punctuation token">(</span>cssCache<span class="punctuation token">.</span><span class="function token">get</span><span class="punctuation token">(</span>from<span class="punctuation token">)</span><span class="punctuation token">)</span><span class="punctuation token">;</span> 407 <span class="keyword token">return</span><span class="punctuation token">;</span> 408 <span class="punctuation token">}</span> 409 <span class="keyword token">let</span> css <span class="operator token">=</span> <span class="string token">''</span><span class="punctuation token">;</span> 410 <span class="keyword token">for</span> <span class="punctuation token">(</span><span class="keyword token">const</span> ruleset <span class="keyword token">of</span> from<span class="punctuation token">)</span> <span class="punctuation token">{</span> 411 <span class="keyword token">let</span> stringifiedDecl <span class="operator token">=</span> <span class="string token">''</span><span class="punctuation token">;</span> 412 <span class="keyword token">for</span> <span class="punctuation token">(</span><span class="keyword token">const</span> <span class="punctuation token">[</span>selector<span class="punctuation token">,</span> declarations<span class="punctuation token">]</span> <span class="keyword token">of</span> Object<span class="punctuation token">.</span><span class="function token">entries</span><span class="punctuation token">(</span> 413 ruleset<span class="punctuation token">,</span> 414 <span class="punctuation token">)</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 415 <span class="keyword token">for</span> <span class="punctuation token">(</span><span class="keyword token">const</span> <span class="punctuation token">[</span>prop<span class="punctuation token">,</span> value<span class="punctuation token">]</span> <span class="keyword token">of</span> Object<span class="punctuation token">.</span><span class="function token">entries</span><span class="punctuation token">(</span> 416 declarations<span class="punctuation token">,</span> 417 <span class="punctuation token">)</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 418 stringifiedDecl <span class="operator token">=</span> <span class="template-string token"><span class="string template-punctuation token">`</span><span class="interpolation token"><span class="interpolation-punctuation punctuation token">${</span>stringifiedDecl<span class="interpolation-punctuation punctuation token">}</span></span><span class="interpolation token"><span class="interpolation-punctuation punctuation token">${</span>prop<span class="interpolation-punctuation punctuation token">}</span></span><span class="string token">:</span><span class="interpolation token"><span class="interpolation-punctuation punctuation token">${</span>value<span class="interpolation-punctuation punctuation token">}</span></span><span class="string token">;</span><span class="string template-punctuation token">`</span></span><span class="punctuation token">;</span> 419 <span class="punctuation token">}</span> 420 css <span class="operator token">=</span> <span class="template-string token"><span class="string template-punctuation token">`</span><span class="interpolation token"><span class="interpolation-punctuation punctuation token">${</span>css<span class="interpolation-punctuation punctuation token">}</span></span><span class="interpolation token"><span class="interpolation-punctuation punctuation token">${</span>selector<span class="interpolation-punctuation punctuation token">}</span></span><span class="string token">{</span><span class="interpolation token"><span class="interpolation-punctuation punctuation token">${</span>stringifiedDecl<span class="interpolation-punctuation punctuation token">}</span></span><span class="string token">}</span><span class="string template-punctuation token">`</span></span><span class="punctuation token">;</span> 421 <span class="punctuation token">}</span> 422 <span class="punctuation token">}</span> 423 <span class="comment token">// Process your CSS however you like</span> 424 <span class="keyword token">const</span> <span class="punctuation token">{</span> <span class="literal-property property token">css</span><span class="operator token">:</span> processedCss <span class="punctuation token">}</span> <span class="operator token">=</span> <span class="keyword token">yield</span> <span class="function token">processCss</span><span class="punctuation token">(</span>css<span class="punctuation token">)</span><span class="punctuation token">;</span> 425 cssCache<span class="punctuation token">.</span><span class="function token">set</span><span class="punctuation token">(</span>from<span class="punctuation token">,</span> processedCss<span class="punctuation token">)</span><span class="punctuation token">;</span> 426 emitter<span class="punctuation token">.</span><span class="function token">write</span><span class="punctuation token">(</span>processedCss<span class="punctuation token">)</span><span class="punctuation token">;</span> 427 <span class="punctuation token">}</span><span class="punctuation token">,</span> 428 <span class="punctuation token">}</span><span class="punctuation token">;</span> 429 <span class="punctuation token">}</span><span class="punctuation token">)</span><span class="punctuation token">;</span></code></pre><p>There, not much code and now adding a font to the site is super easy!</p><p>Eleventy doesn’t make you do everything though, they provide plugins for <a href="https://www.11ty.dev/docs/plugins/image/" target="_blank" rel="noopener">optimising images</a>, <a href="https://www.11ty.dev/docs/plugins/bundle/" target="_blank" rel="noopener">per-page CSS/JS/HTML bundling</a>, <a href="https://www.11ty.dev/docs/plugins/syntaxhighlight/" target="_blank" rel="noopener">syntax highlighting</a> and <a href="https://www.11ty.dev/docs/plugins/official/" target="_blank" rel="noopener">many, many more</a>. There’s also a very <a href="https://11tybundle.dev/" target="_blank" rel="noopener">active community</a> out there full of tips, tricks and guides that’ll help you get the most out of the software.</p><h2 id="leveraging-our-design-system" tabindex="-1">Leveraging our design system</h2><p>The design system we use in the Fastmail client is called Elemental. It’s well suited to displaying content in an information-dense environment without overwhelming users. While in the exploratory phase of rebuilding the site, we quickly discovered our existing spacing and typography didn’t fit our vision for the new site. Everything was too compact! We wanted a new set of <a href="https://thedesignsystem.guide/design-tokens" target="_blank" rel="noopener">design tokens</a> that adapted well to any screen size and reduced the burden on a designer to create different layouts for different viewports.</p><p><a href="https://utopia.fyi/" target="_blank" rel="noopener">Utopia</a> met these needs precisely for us and provided a great framework for thinking about new layouts. In hindsight, I think we could have gotten away with a much simpler set of tokens, but I won’t deny the confidence it gave us in moving forward on this project.</p><h2 id="partial-site-building" tabindex="-1">Partial site building</h2><p>We’ve got quite a few images on our site, mainly due to the number of blog posts we have—<a href="https://www.fastmail.com/blog/imported-older-news-10/" target="_blank" rel="noopener">it goes all the way back to 2001</a>! An unfortunate side-effect of this is that it significantly increases how long a build takes simply due to the sheer number of images we need to process and generate.</p><p>For draft posts, we don’t need to build every post on the blog. In fact, we only need the post (or posts) the author is working on. To achieve this, we can use permalink objects to tell Eleventy that it shouldn’t render the content of a particular page.</p><pre class="language-jsx"><code class="language-jsx"><span class="comment token">// -------------------------</span> 430 <span class="comment token">// source/_data/NO_RENDER.js</span> 431 <span class="comment token">// -------------------------</span> 432 <span class="comment token">// If we set the value of a page's permalink to this object, then the content</span> 433 <span class="comment token">// won't be rendered during build. Under the hood, Eleventy is checking if the</span> 434 <span class="comment token">// object has a `build` property to decide if it should render the page during</span> 435 <span class="comment token">// static generation.</span> 436 <span class="keyword token">export</span> <span class="keyword token">default</span> <span class="punctuation token">{</span> <span class="literal-property property token">norender</span><span class="operator token">:</span> <span class="string token">''</span> <span class="punctuation token">}</span><span class="punctuation token">;</span> 437 438 <span class="comment token">// --------------------------------------</span> 439 <span class="comment token">// source/content/posts/+data.11tydata.js</span> 440 <span class="comment token">// --------------------------------------</span> 441 <span class="comment token">// Now, we can do something like this in a directory data file:</span> 442 <span class="keyword token">function</span> <span class="function token">shouldRenderPost</span><span class="punctuation token">(</span><span class="parameter token">data</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 443 <span class="comment token">// Posts to render</span> 444 <span class="keyword token">return</span> data<span class="punctuation token">.</span>env<span class="punctuation token">.</span>postsToRender 445 <span class="operator token">?</span> data<span class="punctuation token">.</span>env<span class="punctuation token">.</span>postsToRender<span class="punctuation token">.</span><span class="function token">some</span><span class="punctuation token">(</span> 446 <span class="punctuation token">(</span><span class="parameter token">id</span><span class="punctuation token">)</span> <span class="operator token">=></span> 447 id <span class="operator token">===</span> data<span class="punctuation token">.</span>id <span class="operator token">+</span> <span class="string token">''</span> <span class="operator token">||</span> 448 id <span class="operator token">===</span> data<span class="punctuation token">.</span>permalink <span class="operator token">||</span> 449 <span class="template-string token"><span class="string template-punctuation token">`</span><span class="string token">/blog/</span><span class="interpolation token"><span class="interpolation-punctuation punctuation token">${</span>id<span class="interpolation-punctuation punctuation token">}</span></span><span class="string token">/</span><span class="string template-punctuation token">`</span></span> <span class="operator token">===</span> data<span class="punctuation token">.</span>permalink <span class="operator token">||</span> 450 path<span class="punctuation token">.</span><span class="function token">normalize</span><span class="punctuation token">(</span>id<span class="punctuation token">)</span> <span class="operator token">===</span> path<span class="punctuation token">.</span><span class="function token">normalize</span><span class="punctuation token">(</span>data<span class="punctuation token">.</span>page<span class="punctuation token">.</span>inputPath<span class="punctuation token">)</span> 451 <span class="punctuation token">)</span> 452 <span class="operator token">:</span> <span class="boolean token">true</span><span class="punctuation token">;</span> 453 <span class="punctuation token">}</span> 454 455 <span class="keyword token">export</span> <span class="keyword token">default</span> <span class="punctuation token">{</span> 456 <span class="literal-property property token">layout</span><span class="operator token">:</span> <span class="string token">'layouts/post'</span><span class="punctuation token">,</span> 457 <span class="literal-property property token">tags</span><span class="operator token">:</span> <span class="punctuation token">[</span><span class="string token">'posts'</span><span class="punctuation token">]</span><span class="punctuation token">,</span> 458 <span class="literal-property property token">eleventyComputed</span><span class="operator token">:</span> <span class="punctuation token">{</span> 459 <span class="function function-variable token">permalink</span><span class="operator token">:</span> <span class="keyword token">function</span> <span class="punctuation token">(</span><span class="parameter token">data</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 460 <span class="keyword token">return</span> <span class="function token">shouldRenderPost</span><span class="punctuation token">(</span>data<span class="punctuation token">)</span> 461 <span class="operator token">?</span> data<span class="punctuation token">.</span>permalink 462 <span class="operator token">:</span> data<span class="punctuation token">.</span><span class="constant token">NO_RENDER</span><span class="punctuation token">;</span> 463 <span class="punctuation token">}</span><span class="punctuation token">,</span> 464 <span class="punctuation token">}</span><span class="punctuation token">,</span> 465 <span class="punctuation token">}</span><span class="punctuation token">;</span></code></pre><p>You can get the list of posts that have changed with a few git commands. Some of this will depend on your build environment, the following code is applicable to Cloudflare Pages:</p><pre class="language-bash"><code class="language-bash"><span class="important shebang token">#!/usr/bin/env bash</span> 466 467 <span class="builtin class-name token">set</span> <span class="parameter token variable">-euxo</span> pipefail 468 469 <span class="function function-name token">filter</span><span class="punctuation token">(</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 470 <span class="keyword token">while</span> <span class="builtin class-name token">read</span> line<span class="punctuation token">;</span> <span class="keyword token">do</span> 471 <span class="keyword token">for</span> <span class="for-or-select token variable">x</span> <span class="keyword token">in</span> <span class="token variable">$line</span><span class="punctuation token">;</span> <span class="keyword token">do</span> 472 <span class="keyword token">if</span> <span class="token variable">$1</span> <span class="string token">"<span class="token variable">$x</span>"</span><span class="punctuation token">;</span> <span class="keyword token">then</span> 473 <span class="builtin class-name token">echo</span> <span class="string token">"<span class="token variable">$x</span>"</span> 474 <span class="keyword token">fi</span> 475 <span class="keyword token">done</span> 476 <span class="keyword token">done</span> 477 <span class="punctuation token">}</span> 478 479 <span class="function function-name token">isBlogContent</span><span class="punctuation token">(</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 480 <span class="builtin class-name token">local</span> <span class="assign-left token variable">EXT</span><span class="operator token">=</span><span class="token variable">${1<span class="operator token">##</span>*.}</span> 481 <span class="punctuation token">[</span><span class="punctuation token">[</span> <span class="token variable">$1</span> <span class="operator token">==</span> source/content/posts/* <span class="operator token">&amp;&amp;</span> <span class="token variable">$EXT</span> <span class="operator token">==</span> <span class="string token">'md'</span> <span class="punctuation token">]</span><span class="punctuation token">]</span> 482 <span class="punctuation token">}</span> 483 484 <span class="keyword token">if</span> <span class="punctuation token">[</span><span class="punctuation token">[</span> <span class="parameter token variable">-n</span> <span class="string token">"<span class="token variable">${CF_PAGES_BRANCH+x}</span>"</span> <span class="punctuation token">]</span><span class="punctuation token">]</span> <span class="operator token">&amp;&amp;</span> <span class="punctuation token">[</span><span class="punctuation token">[</span> <span class="token variable">$CF_PAGES_BRANCH</span> <span class="operator token">==</span> draft* <span class="punctuation token">]</span><span class="punctuation token">]</span><span class="punctuation token">;</span> <span class="keyword token">then</span> 485 <span class="function token">git</span> remote set-branches origin <span class="string token">'*'</span> 486 <span class="function token">git</span> fetch <span class="parameter token variable">--depth</span><span class="operator token">=</span><span class="number token">100</span> origin production <span class="token variable">$CF_PAGES_BRANCH</span> 487 <span class="assign-left token variable">CHANGES</span><span class="operator token">=</span><span class="token variable"><span class="token variable">$(</span><span class="function token">git</span> <span class="function token">diff</span> --name-only origin/production<span class="punctuation token">..</span>.HEAD <span class="operator token">|</span> filter isBlogContent <span class="operator token">|</span> <span class="function token">tr</span> <span class="string token">'\n'</span> <span class="string token">' '</span> <span class="operator token">|</span> <span class="function token">xargs</span><span class="token variable">)</span></span> 488 <span class="assign-left token variable">POST_IDS</span><span class="operator token">=</span><span class="token variable">$CHANGES</span> npx eleventy 489 <span class="keyword token">else</span> 490 npx eleventy 491 <span class="function token">npm</span> run pagefind:index 492 <span class="keyword token">fi</span></code></pre><p>This significantly sped up build times on draft branches, going from ~5 minutes to ~1:30 per build.</p><h2 id="creating-a-workflow-for-developers" tabindex="-1">Creating a workflow for developers</h2><p>We want people to be able to publish blog posts by simply merging a PR. But if anyone can contribute, how will they know what should go in each field of the front matter? What tags can they use? How do they find their author ID? Sure, you can write some documentation, but wouldn’t it be better to write a few simple scripts that help us enforce content rules?</p><p>These scripts go a long way to empowering a developer to contribute without getting bogged down in the idiosyncrasies of our particular system. Tools like <a href="https://github.com/enquirer/enquirer" target="_blank" rel="noopener">enquirer</a> help you bash out a quick script that keeps your content consistent while providing an easy-to-use CLI.</p><p>Eleventy also provides hooks to <a href="https://www.11ty.dev/docs/data-validate/" target="_blank" rel="noopener">validate your data</a> during build. And don’t forget, if you want to use existing site data in these scripts, use the same tools as Eleventy such as LiquidJS and gray-matter.</p><h2 id="enabling-non-technical-contributors" tabindex="-1">Enabling non-technical contributors</h2><p>So far this is all sounding great if you know how to use git. Just write some markdown and go. But if you’re unfamiliar with git, then writing a simple blog post can seem daunting. We needed a CMS that integrated with our repository, didn’t keep our data behind a proprietary API, and allowed technical and non-technical contributors to use the workflow that felt most comfortable to them.</p><p><a href="https://cloudcannon.com/" target="_blank" rel="noopener">CloudCannon</a> meets each of these requirements and fully enables anyone to contribute to the website. Rather than storing your data, CloudCannon integrates with your site by reading your existing content and providing an interface to create, edit, and delete it. You can further enhance your integration by using their component workflow, <a href="https://github.com/CloudCannon/bookshop" target="_blank" rel="noopener">Bookshop</a>. This will also allow you to <a href="https://cloudcannon.com/documentation/guides/bookshop-eleventy-guide/page-building/" target="_blank" rel="noopener">build pages</a> using a visual editor by composing Bookshop components in their UI.</p><p>I recommend checking out CloudCannon’s <a href="https://github.com/CloudCannon" target="_blank" rel="noopener">GitHub profile</a>, they have a lot of interesting projects geared toward enhancing static sites.</p><h2 id="edge-worker-enhancements" tabindex="-1">Edge worker enhancements</h2><p>We can’t quite get all the way with static assets alone. For our use case, we wanted to display localised currencies on our pricing page based on the user’s country. We chose <a href="https://pages.cloudflare.com/" target="_blank" rel="noopener">Cloudflare Pages</a> as our hosting platform, so here we’ll use a <a href="https://developers.cloudflare.com/pages/functions/" target="_blank" rel="noopener">Pages Function</a> to achieve the desired outcome.</p><p>Cloudflare provide a great interface for transforming HTML on the fly. Simply build an <code>HTMLRewriter</code> instance and use the <code>transform</code> method to modify the body of your response. For us, it looks something like this:</p><pre class="language-jsx"><code class="language-jsx"><span class="comment token">// ------------------------------</span> 493 <span class="comment token">// functions/pricing/[country].ts</span> 494 <span class="comment token">// ------------------------------</span> 495 496 <span class="comment token">// Imported from our 11ty project</span> 497 <span class="keyword token">import</span> billingCountries <span class="keyword token">from</span> <span class="string token">'../../source/_data/billingCountries.json'</span><span class="punctuation token">;</span> 498 499 <span class="keyword token">export</span> <span class="keyword token">const</span> <span class="literal-property property token">onRequest</span><span class="operator token">:</span> PagesFunction<span class="operator token">&lt;</span>Env<span class="punctuation token">,</span> <span class="string token">'country'</span><span class="operator token">></span> <span class="operator token">=</span> <span class="keyword token">async</span> <span class="punctuation token">(</span><span class="parameter token">context</span><span class="punctuation token">)</span> <span class="operator token">=></span> <span class="punctuation token">{</span> 500 <span class="comment token">// ...</span> 501 <span class="keyword token">return</span> <span class="keyword token">new</span> <span class="class-name token">HTMLRewriter</span><span class="punctuation token">(</span><span class="punctuation token">)</span> 502 <span class="punctuation token">.</span><span class="function token">on</span><span class="punctuation token">(</span><span class="string token">'[data-plan-id] [data-swap]'</span><span class="punctuation token">,</span> <span class="punctuation token">{</span> 503 <span class="function token">element</span><span class="punctuation token">(</span><span class="parameter token">el</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 504 el<span class="punctuation token">.</span><span class="function token">replace</span><span class="punctuation token">(</span>swapElements<span class="punctuation token">[</span>el<span class="punctuation token">.</span><span class="function token">getAttribute</span><span class="punctuation token">(</span><span class="string token">'data-swap'</span><span class="punctuation token">)</span><span class="punctuation token">]</span><span class="punctuation token">,</span> <span class="punctuation token">{</span> 505 <span class="literal-property property token">html</span><span class="operator token">:</span> <span class="boolean token">true</span><span class="punctuation token">,</span> 506 <span class="punctuation token">}</span><span class="punctuation token">)</span><span class="punctuation token">;</span> 507 <span class="punctuation token">}</span><span class="punctuation token">,</span> 508 <span class="punctuation token">}</span><span class="punctuation token">)</span> 509 <span class="punctuation token">.</span><span class="function token">on</span><span class="punctuation token">(</span><span class="string token">'[name="select-currency"] option'</span><span class="punctuation token">,</span> <span class="punctuation token">{</span> 510 <span class="function token">element</span><span class="punctuation token">(</span><span class="parameter token">el</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 511 <span class="keyword token">if</span> <span class="punctuation token">(</span>el<span class="punctuation token">.</span><span class="function token">hasAttribute</span><span class="punctuation token">(</span><span class="string token">'selected'</span><span class="punctuation token">)</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 512 el<span class="punctuation token">.</span><span class="function token">removeAttribute</span><span class="punctuation token">(</span><span class="string token">'selected'</span><span class="punctuation token">)</span><span class="punctuation token">;</span> 513 <span class="punctuation token">}</span> 514 <span class="keyword token">if</span> <span class="punctuation token">(</span>el<span class="punctuation token">.</span><span class="function token">getAttribute</span><span class="punctuation token">(</span><span class="string token">'value'</span><span class="punctuation token">)</span><span class="punctuation token">.</span><span class="function token">toLowerCase</span><span class="punctuation token">(</span><span class="punctuation token">)</span> <span class="operator token">===</span> country<span class="punctuation token">)</span> <span class="punctuation token">{</span> 515 el<span class="punctuation token">.</span><span class="function token">setAttribute</span><span class="punctuation token">(</span><span class="string token">'selected'</span><span class="punctuation token">,</span> <span class="string token">''</span><span class="punctuation token">)</span><span class="punctuation token">;</span> 516 <span class="punctuation token">}</span> 517 <span class="punctuation token">}</span><span class="punctuation token">,</span> 518 <span class="punctuation token">}</span><span class="punctuation token">)</span> 519 <span class="punctuation token">.</span><span class="function token">on</span><span class="punctuation token">(</span><span class="string token">'[data-tax-notice]'</span><span class="punctuation token">,</span> <span class="punctuation token">{</span> 520 <span class="function token">element</span><span class="punctuation token">(</span><span class="parameter token">el</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 521 el<span class="punctuation token">.</span><span class="function token">setInnerContent</span><span class="punctuation token">(</span>swapElements<span class="punctuation token">[</span><span class="string token">'taxNotice'</span><span class="punctuation token">]</span><span class="punctuation token">)</span><span class="punctuation token">;</span> 522 <span class="punctuation token">}</span><span class="punctuation token">,</span> 523 <span class="punctuation token">}</span><span class="punctuation token">)</span> 524 <span class="punctuation token">.</span><span class="function token">on</span><span class="punctuation token">(</span><span class="string token">'[name="plan-type"]'</span><span class="punctuation token">,</span> <span class="punctuation token">{</span> 525 <span class="function token">element</span><span class="punctuation token">(</span><span class="parameter token">el</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 526 <span class="keyword token">if</span> <span class="punctuation token">(</span><span class="operator token">!</span>url<span class="punctuation token">.</span>searchParams<span class="punctuation token">.</span><span class="function token">has</span><span class="punctuation token">(</span><span class="string token">'plan-type'</span><span class="punctuation token">,</span> <span class="string token">'business'</span><span class="punctuation token">)</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 527 <span class="keyword token">return</span><span class="punctuation token">;</span> 528 <span class="punctuation token">}</span> 529 <span class="keyword token">const</span> isBusiness <span class="operator token">=</span> el<span class="punctuation token">.</span><span class="function token">getAttribute</span><span class="punctuation token">(</span><span class="string token">'value'</span><span class="punctuation token">)</span> <span class="operator token">===</span> <span class="string token">'business'</span><span class="punctuation token">;</span> 530 <span class="keyword token">if</span> <span class="punctuation token">(</span>isBusiness<span class="punctuation token">)</span> <span class="punctuation token">{</span> 531 el<span class="punctuation token">.</span><span class="function token">setAttribute</span><span class="punctuation token">(</span><span class="string token">'checked'</span><span class="punctuation token">,</span> <span class="string token">'checked'</span><span class="punctuation token">)</span><span class="punctuation token">;</span> 532 <span class="punctuation token">}</span> <span class="keyword token">else</span> <span class="punctuation token">{</span> 533 el<span class="punctuation token">.</span><span class="function token">removeAttribute</span><span class="punctuation token">(</span><span class="string token">'checked'</span><span class="punctuation token">)</span><span class="punctuation token">;</span> 534 <span class="punctuation token">}</span> 535 <span class="punctuation token">}</span><span class="punctuation token">,</span> 536 <span class="punctuation token">}</span><span class="punctuation token">)</span> 537 <span class="punctuation token">.</span><span class="function token">transform</span><span class="punctuation token">(</span><span class="keyword token">await</span> context<span class="punctuation token">.</span>env<span class="punctuation token">.</span><span class="constant token">ASSETS</span><span class="punctuation token">.</span><span class="function token">fetch</span><span class="punctuation token">(</span>pageURL<span class="punctuation token">)</span><span class="punctuation token">)</span><span class="punctuation token">;</span> 538 <span class="punctuation token">}</span> 539 540 <span class="comment token">// --------------------------</span> 541 <span class="comment token">// functions/pricing/index.ts</span> 542 <span class="comment token">// --------------------------</span> 543 544 <span class="keyword token">import</span> billingCountries <span class="keyword token">from</span> <span class="string token">'../../source/_data/billingCountries.json'</span><span class="punctuation token">;</span> 545 546 <span class="keyword token">const</span> supportedCountries <span class="operator token">=</span> billingCountries<span class="punctuation token">.</span><span class="function token">reduce</span><span class="punctuation token">(</span><span class="punctuation token">(</span><span class="parameter token">countries<span class="punctuation token">,</span> country</span><span class="punctuation token">)</span> <span class="operator token">=></span> <span class="punctuation token">{</span> 547 countries<span class="punctuation token">.</span><span class="function token">add</span><span class="punctuation token">(</span>country<span class="punctuation token">.</span>code<span class="punctuation token">)</span><span class="punctuation token">;</span> 548 <span class="keyword token">return</span> countries<span class="punctuation token">;</span> 549 <span class="punctuation token">}</span><span class="punctuation token">,</span> <span class="keyword token">new</span> <span class="class-name token">Set</span><span class="punctuation token">(</span><span class="punctuation token">)</span><span class="punctuation token">)</span><span class="punctuation token">;</span> 550 551 <span class="keyword token">export</span> <span class="keyword token">const</span> <span class="literal-property property token">onRequest</span><span class="operator token">:</span> PagesFunction<span class="tag token"><span class="tag token"><span class="punctuation token">&lt;</span><span class="class-name token">Env</span></span><span class="punctuation token">></span></span><span class="plain-text token"> = (context) => </span><span class="punctuation token">{</span> 552 <span class="keyword token">let</span> country <span class="operator token">=</span> context<span class="punctuation token">.</span>request<span class="punctuation token">.</span>cf<span class="punctuation token">.</span>country<span class="punctuation token">;</span> 553 <span class="keyword token">if</span> <span class="punctuation token">(</span><span class="operator token">!</span>supportedCountries<span class="punctuation token">.</span><span class="function token">has</span><span class="punctuation token">(</span>country<span class="punctuation token">)</span><span class="punctuation token">)</span> <span class="punctuation token">{</span> 554 country <span class="operator token">=</span> <span class="string token">'US'</span><span class="punctuation token">;</span> 555 <span class="punctuation token">}</span> 556 <span class="keyword token">const</span> url <span class="operator token">=</span> <span class="keyword token">new</span> <span class="class-name token">URL</span><span class="punctuation token">(</span>context<span class="punctuation token">.</span>request<span class="punctuation token">.</span>url<span class="punctuation token">)</span><span class="punctuation token">;</span> 557 url<span class="punctuation token">.</span>pathname <span class="operator token">=</span> <span class="template-string token"><span class="string template-punctuation token">`</span><span class="string token">/pricing/</span><span class="interpolation token"><span class="interpolation-punctuation punctuation token">${</span>country<span class="punctuation token">.</span><span class="function token">toLowerCase</span><span class="punctuation token">(</span><span class="punctuation token">)</span><span class="interpolation-punctuation punctuation token">}</span></span><span class="string token">/</span><span class="string template-punctuation token">`</span></span><span class="punctuation token">;</span> 558 <span class="keyword token">return</span> Response<span class="punctuation token">.</span><span class="function token">redirect</span><span class="punctuation token">(</span>url<span class="punctuation token">.</span>href<span class="punctuation token">,</span> <span class="number token">302</span><span class="punctuation token">)</span><span class="punctuation token">;</span> 559 <span class="punctuation token">}</span><span class="plain-text token">; 560 </span></code></pre><p>Now, when a user visits <a href="http://www.fastmail.com/pricing/" target="_blank" rel="noopener">fastmail.com/pricing/</a>, they’ll be redirected to a page with prices in their currency (if we support it).</p><p>Unfortunately, this is the part of our site that isn’t portable. Hopefully, once <a href="https://wintercg.org/" target="_blank" rel="noopener">WinterCG</a> picks up steam, we’ll start seeing much better compatibility between these different runtimes. But for now, using an Edge worker does mean accepting a certain level of lock-in with your provider.</p><h2 id="wrapping-up" tabindex="-1">Wrapping up</h2><p>Personally, I’ve found Eleventy to be a great piece of software. It puts you in full control of the site you want to build and gives you the tools to extend its functionality with ease.</p></content> 561 </entry><entry> 562 <title>Dec 8: Guiding principles</title> 563 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/principles/' /> 564 <id>https://www.fastmail.com/blog/principles/</id> 565 <updated>2024-12-08T00:00:01Z</updated><author> 566 <name>Bron Gondwana</name> 567 </author><content xml:lang='en' type='html'><p>This is the eighth post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/revision-of-core-email-specifications/">Dec 7: Revision of the core email specifications</a>. The next post is <a href="/blog/building-a-blog/">Dec 9: Building a blog</a>.</p><p>When it’s a week before December and you suddenly remember “oh yeah, we were going to do an advent blog series” — it does help to have some pre-written material! So I’m cheating a bit and pulling from our internal Notion, getting the words we’ve already spent a bunch of time honing.</p><p>I wrote last Sunday about our <a href="/blog/mission-statement/">Mission Statement</a>. A mission statement is all well and good, but it’s very high level and aspirational. It needs to be distilled down to something that’s actionable and useful on a daily basis.</p><p>So we wrote the following:</p><h2 id="our-guiding-principles" tabindex="-1">Our guiding principles</h2><p>How we go about achieving that vision is also very important to who Fastmail is as a company. We feel these principles really capture how Fastmail has operated as a company in the past, and should do in the future.</p><blockquote> <h3 id="we-are-a-good-internet-citizen" tabindex="-1">We are a good internet citizen</h3> <ul> <li>We believe in open protocols, standards and interoperability.</li> <li>We build, share, and support technology to make email better for everyone, not another walled garden.</li> <li>We foster positive relationships with our customers, partners, suppliers, and staff.</li> </ul> </blockquote><p>This is also one of our <a href="/company/values/">public values</a>. We encourage our employees to contribute to open source projects, to be involved in the technical communities in their local area, and to be involved in the standards development process.</p><blockquote> <h3 id="we-build-the-future" tabindex="-1">We build the future</h3> <ul> <li>We are not content to just accept the status quo.</li> <li>If the right tool doesn’t exist, we make it.</li> <li>If the open standards aren’t good enough, we improve them.</li> <li>We are leaders in our industry.</li> </ul> </blockquote><p>Yes, we’re a bit “not invented here”. We maintain the <a href="https://cyrusimap.org/" target="_blank" rel="noopener">Cyrus IMAP</a> server. We built our own <a href="https://github.com/fastmail/overture" target="_blank" rel="noopener">Javascript framework</a> and <a href="https://github.com/fastmail/Squire" target="_blank" rel="noopener">email editor</a>. We buy our own hardware made to spec for what we need, and manage everything from the operating system up.</p><p>We like to make things ourselves, because then we understand how they work; which leads into the next point!</p><blockquote> <h3 id="we-seek-understanding" tabindex="-1">We seek understanding</h3> <ul> <li>We are curious about how things work, and how they came to be.</li> <li>We seek deep and actionable understanding of our systems.</li> <li>We recognise that we can’t understand everything, but strive to know where the boundaries of our knowledge lie.</li> </ul> </blockquote><p>We really don’t like unexplained behaviours in our system. If we don’t know why it happened, that’s a problem! It’s this attitude which led us (see the past point) to <a href="/blog/twoskip-and-more/">debug and then replace the skiplist database format in Cyrus</a>.</p><p>We’re not writing our own filesystem or operating system though! At least not this week.</p><blockquote> <h3 id="we-value-discussion" tabindex="-1">We value discussion</h3> <ul> <li>We reach agreement through constructive, iterative collaboration.</li> <li>Everyone is encouraged to express their theories, and explain the basis on which they are formed.</li> <li>We test our assumptions against the real world.</li> <li>What works is more important than who thought of it.</li> </ul> </blockquote><p>We really believe in confirming theories against reality. Nothing beats a good hypothesis, and a well constructed experiment to determine whether it adequately predicted what happens.</p><hr><p>These principles really do reflect how we think and talk about ourselves at Fastmail. We live by these every day.</p><p>There have been times when we have considered that we need an additional principle “We get shit done”, which is in a degree of opposition to both “seek understanding” and “value discussion”. That’s the thing with principles like this, they are a tradeoff between different valuable properties. These are the ones we choose to prioritize! So we discuss, we understand, then we do.</p></content> 568 </entry><entry> 569 <title>Dec 7: Revision of the core email specifications</title> 570 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/revision-of-core-email-specifications/' /> 571 <id>https://www.fastmail.com/blog/revision-of-core-email-specifications/</id> 572 <updated>2024-12-07T00:00:01Z</updated><author> 573 <name>Ken Murchison</name> 574 </author><content xml:lang='en' type='html'><p>This is the seventh post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/twoskip-and-more/">Dec 6: Twoskip and more</a>. The next post is <a href="/blog/principles/">Dec 8: Guiding principles</a>.</p><p>There are many specifications that form the current internet messaging ecosystem, but the most integral of these are <a href="https://datatracker.ietf.org/doc/rfc5321/" target="_blank" rel="noopener">RFC 5321</a> and <a href="https://datatracker.ietf.org/doc/rfc5322/" target="_blank" rel="noopener">RFC 5322</a>. RFC 5321 specifies the protocol for the transport of electronic mail messages on the internet, known as the Simple Mail Transfer Protocol (SMTP). RFC 5322 specifies the base syntax of these messages, known as the Internet Message Format (IMF). These documents were both published in 2008, and each are the third version of the “core” specifications (see RFCs <a href="https://datatracker.ietf.org/doc/rfc821/" target="_blank" rel="noopener">821</a> / <a href="https://datatracker.ietf.org/doc/rfc2821/" target="_blank" rel="noopener">2821</a> and RFCs <a href="https://datatracker.ietf.org/doc/rfc822/" target="_blank" rel="noopener">822</a> / <a href="https://datatracker.ietf.org/doc/rfc2822/" target="_blank" rel="noopener">2822</a> respectively).</p><p>Over the sixteen years since publication, several errata have been filed against both documents. As a result, an effort is currently underway within the Internet Engineering Task Force (IETF) <a href="https://datatracker.ietf.org/wg/emailcore/about/" target="_blank" rel="noopener">EMAILCORE</a> working group to revise these specifications yet again. Per its charter, the working group is tasked with updating the documents with “corrections and clarifications only, with a strong emphasis on keeping these minimal and avoiding broader changes to terminology or document organization”.</p><p>Work on the revision to the transport protocol specification (<a href="https://datatracker.ietf.org/doc/draft-ietf-emailcore-rfc5321bis/" target="_blank" rel="noopener">RFC 5321bis</a>) continues on a few remaining issues, but is nearing completion. Work on the revision to the message format specification (<a href="https://datatracker.ietf.org/doc/draft-ietf-emailcore-rfc5322bis/" target="_blank" rel="noopener">RFC 5322bis</a>) has already been completed, provided that no further changes are prompted by changes to RFC 5321bis. Remarkably, both documents have the same editors as the previous two revisions going back to 2001 — John Klensin and Pete Resnick respectively. For those that are interested in what has been corrected and/or clarified in these revisions, each of the documents have appendices that contain a list of changes, with the discussions of these changes having taken place on the working group’s <a href="https://mailarchive.ietf.org/arch/browse/emailcore/" target="_blank" rel="noopener">mailing list</a>.</p><p>In addition to revising the core specifications, the group is also working on an <a href="https://datatracker.ietf.org/doc/draft-ietf-emailcore-as/" target="_blank" rel="noopener">applicability statement</a> to document other relevant specifications that implementors should be aware of, such as the use of Multipurpose Internet Mail Extensions (MIME) and Transport Layer Security (TLS). It also documents current email best practices, such as which provisions of the transport protocol and message format have proven to cause interoperability issues, and how to properly reuse an existing email message as a template for a new one.</p><p>All three documents are expected to be submitted to the RFC Editor for publication in early 2025.</p></content> 575 </entry><entry> 576 <title>Dec 6: Twoskip and more</title> 577 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/twoskip-and-more/' /> 578 <id>https://www.fastmail.com/blog/twoskip-and-more/</id> 579 <updated>2024-12-06T00:00:01Z</updated><author> 580 <name>Bron Gondwana</name> 581 </author><content xml:lang='en' type='html'><p>This is the sixth post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post is <a href="/blog/mysql-innodb-trx-id/">Dec 5: MySQL InnoDB innodb_trx is cached</a>. The next post is <a href="/blog/revision-of-core-email-specifications/">Dec 7: Revision of the core email specifications</a>.</p><p>Rob wrote yesterday about some of the issues he ran into using MySQL triggers to implement JMAP access to our database tables, so I figured I’d follow up with talking about the other database that a lot of our data is stored in.</p><p>I love talking about twoskip. It’s used within the <a href="https://cyrusimap.org/" target="_blank" rel="noopener">Cyrus IMAP server</a> for all our big databases. And I wrote something about <a href="/blog/cyrus-databases-twoskip-and-beyond/">possible futures for the format</a> back in the 2016 advent. Well, we over-engineered our attempt at zeroskip, a more copy-on-write friendly alternative to twoskip, so we’ve been running twoskip basically unmodified since then despite moving to ZFS for all our email storage.</p><p>It turns out that when you have really fast NVMe, even suboptimal filesystem usage is still pretty fast.</p><h2 id="before-twoskip" tabindex="-1">Before twoskip</h2><p>When I first started at Fastmail, we were using the Cyrus with the skiplist database format. In the early days, it could get corrupted pretty easily on a busy server. I wrote a tool in perl back in 2006:</p><pre><code>commit 20e5db17fd5be878a84c9b122fc22f49fe0f7ea2 582 Author: Bron Gondwana &lt;brong@fastmailteam.com&gt; 583 Date: Wed Sep 27 02:59:54 2006 +0000 584 585 tool to dump even corrupted skiplists in something resembling a usable form 586 587 diff --git a/utils/oneoff/skiplist_dump.pl b/utils/oneoff/skiplist_dump.pl 588 new file mode 100644 589 index 0000000000..e8bf1b27d1 590 --- /dev/null 591 +++ b/utils/oneoff/skiplist_dump.pl 592 @@ -0,0 +1,115 @@ 593 +#!/usr/bin/perl 594 + 595 +use strict; 596 +use warnings; 597 + 598 +use IO::File; 599 +use Data::Dumper; 600 + 601 +use constant INORDER =&gt; 1; 602 +use constant ADD =&gt; 2; 603 +use constant DELETE =&gt; 4; 604 +use constant COMMIT =&gt; 255; 605 +use constant DUMMY =&gt; 257; 606 + 607 +my $fh = IO::File-&gt;new(shift); 608 </code></pre><p>Because we were getting so much corruption, I wound up figuring out what was going on and writing a patch to Cyrus. Here was the description:</p><pre><code>conf/patches/cyrus_quilt/cyrus-skiplist-bugfixes-2.3.10.diff 609 610 SKIPLIST bugfixes 611 612 &lt;b&gt;ACCEPTED UPSTREAM 2007-11-16&lt;/b&gt; 613 614 In the past we have had issues with bugs in skiplist on seen 615 files, and we truncated files at the offset with the issue 616 since they were only seen data. 617 618 Lately, we have had more tools updating mailboxes.db more 619 often, and have lost multiple mailboxes.db files. 620 621 There are two detectable issues: 622 623 1) incorrect header &quot;logstart&quot; values causing recovery to 624 fail with either unexpected ADD/DELETE records or 625 unexpected INORDER records depending which side of the 626 correct location the logstart value is wrong. 627 2) a bunch of zero bytes between transactions in the log 628 section. 629 630 The attached patch fixes the following issues: 631 632 a) recovery failed to update db-&gt;map_base if it truncated 633 a partial transaction. This reliably recreated the 634 zero bytes issues above by having the next store command 635 lseek to a location past the new end of the file, and 636 hence fill the remainder with blanks. 637 638 b) the logic in the &quot;delete&quot; handler for detecting &quot;no 639 record exists&quot; (ptr == db-&gt;map_base) was backwards, 640 meaning that a delete on a record which didn't exist 641 caused reads of PTR(db-&gt;map_base, i), which is bogus 642 and nasty. This is the suspect for logstart breakage 643 though I haven't proven this yet. 644 645 c) unsure if this is a real risk, but added a ftruncate 646 to checkpoint to ensure new file really is empty, 647 since we don't open it with O_EXCL. 648 649 d.1) when abort is called it needs to update_lock() to 650 ensure that the records it's about to rollback are 651 actually locked. This fix stopped segfaults in 652 my testing. 653 654 d.2) delete didn't check retry_write for success, and 655 also suffered from the same problem as: 656 657 d.3) if retry_writev in store failed, then it called 658 abort, but without the records actually written, 659 abort had no way of knowing which offsets needed 660 to be switched back, meaning bogus pointers could 661 be left in the file until the next recovery(). 662 Changed it to write the records first, then only 663 once that succeeded update the pointers. This way 664 abort will do the right thing regardless. 665 666 *PHEW* - that's three days I want back, but it has survived 667 10 concurrent processes doing nasty things to it for 10000 668 operations each and it's still going strong: 669 skiplist: recovered /tmp/hammer.db (143177 records, 5965096 bytes) in 0 seconds 670 </code></pre><p>So that was that. But skiplist was still expensive to recover after a crash, so by 2011 I had replaced it with twoskip. Twoskip had many design goals (this is all in comments at the top of the source file, you can <a href="https://github.com/cyrusimap/cyrus-imapd/tree/master/lib/cyrusdb_twoskip.c" target="_blank" rel="noopener">read it on Github</a>).</p><h3 id="design-goals" tabindex="-1">Design Goals:</h3><ul> <li>64 bit throughout</li> <li>Checksums on all content</li> <li>Single on-disk file</li> <li>Fully embedded, no daemons or external helper</li> <li>Readable without causing writes (if clean)</li> <li>May require that an aligned 512 bytes either fully writes or doesn’t write anything.</li> <li>May require fsync to work</li> </ul><p>We achieved all those, which is why twoskip has been such a workhorse. With that said,</p><h2 id="a-corrupt-twoskip-file" tabindex="-1">A corrupt twoskip file</h2><p>Last year, we had our only incident of corrupt twoskip files in this entire time!</p><p>I’m basically just going to dump the entire Topicbox thread here, because it’s a peek behind the scenes at how we debugged the issue, and also how having redundant servers with failover meant that we could do all this debugging and recovery without impacting access to email for users.</p><h3 id="first-email-by-me" tabindex="-1">First email, by me:</h3><p>Sunday, February 05, 2023 12:28</p><blockquote> <p>I got paged about:</p> <pre><code>2023-02-04T15:54:56.525746-05:00 imap41 sloti41n57/syncserver[2135571]: DBERROR: twoskip checksum head error: filename=&lt;/mnt/i41n57/sloti41n57/store56/conf/user/uuid/c/9/c993a9a5-5fed-4515-807e-0b1326c7200d/conversations.db&gt; offset=&lt;00370B30&gt; syserror=&lt;No such file or directory&gt; func=&lt;read_onerecord&gt; 671 </code></pre> <p>It turns out there’s a block of corruption from 00360000 to 00380000 in that file. And it’s not all zeros. I’m still investigating.</p> <p>Anyway, failed everything off the machine, rebooted - corruption is still there. Initiated a new zfs scrub. That slot is left down, all others are up.</p> <p>Marc and I are both heading out for a bit to do personal things today, but I’ll look in again on the scrub later and also keep looking at the corruption. It looks like this:</p> <pre><code>00360000 2b 01 00 34 00 00 00 19 00 00 00 00 00 76 e3 80 |+..4.........v..| 672 00360010 00 00 00 00 00 36 00 70 28 b1 a3 a4 fa ef 97 5e |.....6.p(......^| 673 00360020 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 |................| 674 * 675 003600e0 2b 01 00 30 00 00 00 19 00 00 00 00 00 82 4a c8 |+..0..........J.| 676 003600f0 00 00 00 00 00 36 01 50 11 20 22 7c 62 dd 0d d3 |.....6.P. &quot;|b...| 677 00360100 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 |................| 678 * 679 00360150 2b 02 00 31 00 00 00 19 00 00 00 00 00 00 00 00 |+..1............| 680 00360160 00 00 00 00 00 36 01 c8 00 00 00 00 00 36 03 f8 |.....6.......6..| 681 00360170 62 6b f2 75 12 c0 98 2e 00 00 00 00 00 00 00 00 |bk.u............| 682 00360180 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 |................| 683 * 684 00360230 00 00 00 00 00 00 00 00 2b 01 00 31 00 00 00 19 |........+..1....| 685 00360240 00 00 00 00 00 75 4d d8 00 00 00 00 00 36 02 a8 |.....uM......6..| 686 00360250 f8 55 a3 e8 24 1c 9e 87 00 00 00 00 00 00 00 00 |.U..$...........| 687 00360260 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 |................| 688 * 689 00360310 00 00 00 00 00 00 00 00 2b 01 00 30 00 00 00 19 |........+..0....| 690 00360320 00 00 00 00 00 76 df 08 00 00 00 00 00 36 03 88 |.....v.......6..| 691 00360330 0d 41 4e 13 67 8d aa b4 00 00 00 00 00 00 00 00 |.AN.g...........| 692 00360340 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 |................| 693 694 </code></pre> <p>Which is so weird because it looks like keys but not values being written or something. But the checksums also aren’t consistent so it’s not pure on-write corruption. I don’t know yet.</p> <p>I’m not thrilled with the number of filesystem level blue moons we’ve had in the past few weeks. I want to do more digging later, but for now we’ll let the scrub run.</p> </blockquote><h3 id="i-came-back-later-with-more-detail" tabindex="-1">I came back later with more detail</h3><p>Sunday, February 05, 2023 22:04</p><blockquote> <p>OK - this IS consistent with a zeroed out area being written over later. So the file was written, then this block got zeroed, then later writes came back and updated just the headers of these records, which is how it works.</p> <p>So it SMELLS like maybe this is a pwrite vs mmap issue. Linux (maybe ZFS) failed to page in the content, and then wrote bits of it, and wrote back the block with zeros for all the bits it wasn’t writing.</p> <p>So it’s for sure a page cache bug. OUCH. We can maybe work around this by using mmap for writes as well as reads, certainly if we see this again we’re in trouble.</p> <p>I looked to see if there were ANY other files of the same type on the same machine, and no - everything else was consistent. This took hours:</p> <pre><code>[fastmail root(brong)@imap41 ~]# find /mnt/*/*/*/conf/user/uuid -name conversations.db | xargs -I{} sh -c 'echo -n {} &quot;: &quot;; sudo -u cyrus /mnt/i41n57/sloti41n57/store56/conf/pkg/bin/cyr_dbtool -C /etc/cyrus/imapd-sloti41n57.conf {} twoskip consistent' | grep No 695 /mnt/i41n57/sloti41n57/store56/conf/user/uuid/c/9/c993a9a5-5fed-4515-807e-0b1326c7200d/conversations.db : No, not consistent 696 [fastmail root(brong)@imap41 ~]# 697 698 </code></pre> <p>So I moved it aside and reconstructed the user:</p> <pre><code>[fastmail root(brong)@imap41 ~]# mv /mnt/i41n57/sloti41n57/store56/conf/user/uuid/c/9/c993a9a5-5fed-4515-807e-0b1326c7200d/conversations.db /mnt/i41n57/sloti41n57/store56/conf/user/uuid/c/9/c993a9a5-5fed-4515-807e-0b1326c7200d/conversations.db.old 699 [fastmail root(brong)@imap41 ~]# cyr sloti41n57 ctl_conversationsdb -R [REDACTED USER] 700 [REDACTED DETAIL] 701 [fastmail root(brong)@imap41 ~]# 702 703 </code></pre> <p>And finally started up the slot.</p> </blockquote><h3 id="rob-m-our-cto-chimed-in" tabindex="-1">Rob M, our CTO, chimed in</h3><p>Monday, February 06, 2023 10:26</p><blockquote> <p>I note that this is 128kb, which is the ZFS recordsize we use for the cyrus spools. That seems awfully coincidental.</p> <p>It’s well known that the way ZFS uses memory (ARC) and it’s interaction with the Linux page cache is messy and caused problems in the past, but that it had been basically dealt with.</p> <p>This problem though suggests there might still be a very rare hole somewhere. Given this is the first corruption we’ve seen in over a year of heavy usage across many machines, it seems it’s going to be a very very rare edge case, and likely to be extremely hard to ever track down :(</p> </blockquote><h3 id="rob-n-senior-sysadmin" tabindex="-1">Rob N, Senior Sysadmin</h3><p>Saturday, February 11, 2023 12:21</p><blockquote> <p>This is likely to be https://github.com/openzfs/zfs/issues/13608, or very adjacent to it. Its describing an mmap read returning wrong/corrupted data.</p> <p>No one yet really has a handle on it. It is being investigated a little elsewhere and was discussed at the OpenZFS dev call in early January. It seems clear that its possible for the page cache to be dirty when ZFS thinks it should be clean, but no one has been able to track it down reliably. It might be related to concurrent mmap reads and modifies over the same record, and then being invalidated inconsistently (which is why clones are in play in the above issue).</p> <p>Assuming it is the same issue, we might in a position to assist because we’re coming at it from a different angle.</p> <p>What I want to do next is try to determine the rough sequence events (mapped read -&gt; write -&gt; read -&gt; whatever) including any other programs that might have been operating on the same data at the same time. I suspect we can put together a rough sequence just by reading code. Bron, if you have some time this week, I would like to go through this with you (mostly: show me the code and I’ll nod along). If we can pin down a clear sequence, maybe we can write something that can reproduce it if we drive it hard enough. And if not, maybe we can at least write this all down for the ZFS folk to consider.</p> </blockquote><h3 id="and-finally-rob-n-wrote-up-a-great-summary" tabindex="-1">And finally Rob N wrote up a great summary</h3><p>Sunday, February 19, 2023 22:08</p><blockquote> <p>Bron and I spent a couple of hours last Tuesday going over exactly how twoskip does its work and trying to get a sense of what happened. This is something of a writeup. Thin on technical detail, because that would take even longer for me to write.</p> <p><strong>Tiny tiny twoskip overview</strong></p> <p>It helps to know what a twoskip file is before we talk about how it writes to disk.</p> <p>From outside, its a KV store, with variable-length keys and values. Internally, its effectively a linked list of records, sorted by key (its a skiplist, so there’s actually multiple linked lists, but that doesn’t matter for our purposes).</p> <p>Each record is either a “RECORD”, that is, a real active item in the store, or a “DELETE”, which is a tombstone in place of a thing that used to exist. (There’s a couple of others, that don’t matter). Making a modification means writing a new record to the end of the file, and then going back through the file to fix the linked list pointers to hook that record up in the right place (if you ever wrote a sorted linked list and remember the fiddling to insert a thing in the middle of the list, its just a slightly fancier version of that).</p> <p>Cyrus does fairly regular repack operations, which means creating a new file and writing all the “live” records to it, removing any RECORDs that have a DELETE, and of course writing all the records in order in the file, rather than having pointers back and forth through the file. Same end result, but a better optimised file.</p> <p><strong>Reading and writing</strong></p> <p>Accessing a twoskip file is done via two different mechanisms: <code>mmap()</code> for reads, <code>write()</code> (actually <code>pwritev()</code>) for writes.</p> <p>For the uninitiated, <code>mmap()</code> is (in this case) a way to pretend a file is really a region of memory. Rather than having to read the contents of a file into memory piece by piece, you can just tell the kernel to give you a big chunk of memory. When you try to read from it, the kernel will go and get the data from the file and make it appear in that memory space, so you can kind of pretend that its already been read into memory. The main reason to do this is performance - the kernel is usually given you a view over its own filesystem cache memory, where its already likely got a copy of a busy file available, and it doesn’t have to copy that data into a user-space memory buffer.</p> <p>For writing, we doing a more conventional <code>write()</code> at specific offsets. You can set up a <code>mmap()</code> for writes too, but for small random writes throughout a file (as twoskip writes are) it doesn’t really gain much on performance (the kernel has to copy from userspace) and can be tougher for the user program to manage (in a variety of subtle handwavey ways that tbh I am not an expert in). The calls to <code>write()</code> will be reflected in the existing read map anyway (the kernel will copy it there before it sends it down to disk anyway), so it all works out quite nicely in the end.</p> <p><strong>The twoskip write cycle</strong></p> <p>When we open a twoskip file, we maps the entire file into memory with <code>mmap()</code>. The map region is the length of the file, rounded up to the nearest 16K boundary. At this point we could happily read the existing contents out of memory.</p> <p>When we write a new record, we seek to the end of the file, and then write four distinct parts (in a single <code>write()</code> call):</p> <ul> <li>a 32-byte header, which includes the lengths of the key and value, linked list pointers, and checksums for the header and the data</li> <li>variable-length key</li> <li>variable-length value</li> <li>0-7 padding bytes to bring the entire record to an 8-byte boundary</li> </ul> <p>Then, we go back through the file (before the record), finding the nodes that should point to this new record, and overwrite their 32-byte header updating the pointers to splice it into the list. Each header is a separate call to <code>write()</code>.</p> <p>Now, a <code>mmap()</code> region has a fixed length, and we started with it as the length of the file, plus padding. After writing a record, the file may now actually be larger than the map. There’s nothing wrong with this; <code>mmap()</code> is done by offset and length; it doesn’t have to cover the whole file. This does of course mean, if we have just written a record past the end of the map, we can’t actually read it back (it would cause a segfault). To keep everything nice, after each <code>write()</code>, we check to see if we’ve gone past the end of the map and if so, we extend it out past the end (to the nearest 16K).</p> <p>An important thing to note: if you <code>mmap()</code> past the end of a file (and we do as a matter of course, with the 16K alignment) and then you read from that area, you get zeroes back.</p> <p><strong>ZFS is not like other filesystems</strong></p> <p>There’s a few things to know about ZFS that is different to a more conventional overwriting filesystem in this situation.</p> <p>ZFS’ unit of IO is the “record size”. Its configurable, but we’re using the default, which is 128K. Which is to say, any write to an existing file of smaller than 128K requires loading a full 128K record from disk, modifying it in memory, then writing the full 128K back down again (to a new record; ZFS never overwrites existing records, but that’s not really relevant here).</p> <p>ZFS does not use the kernel-provided filesystem cache (the “page cache”) because it has its own cache (the ARC) which has quite different characteristics to support all sorts of ZFS features. However, <code>mmap()</code> by definition exposes memory in the page cache. To make this work, ZFS copies data between the page cache and the ARC as necessary when a file is mapped.</p> <p>For asynchronous writes, a conventional filesystem will end up having written data staged in the page cache, which is periodically flushed to disk. ZFS has the same concept, but it uses the ARC as the cache, and writes are attached to a “transaction group”, which is flushed periodically. This is mostly interesting here because I’m going to say “transaction group” soon. The details past that don’t matter, mostly its just a collection of writes waiting to go to the storage pool.</p> <p><strong>What even happened?</strong></p> <p>So, what we saw is a single 128K section of the file that had record headers, but zeroes through the data sections. Recall that new records are written to the end of the file, making it larger, then we extend the map over it and past the real end, and then we go back through the file and update the headers.</p> <p>This record has headers, but no data. It seems that the only way this could have happened is if the initial record writes were lost somehow, but the header writes survived. The initial record writes are the only thing that can write “outside” of the currently mapped region; the header fixups always happen on existing records that the map will always be covering.</p> <p>This and the fact that the data is all-zeroes (and not scrambled random memory contents) suggests that <code>write()</code> successfully staged stuff to be written (that is, an ARC buffer for the record was allocated and was attached to a transaction group), but when <code>mmap()</code> was called to extend the map, ZFS decided that the page cache contents (all zeroes) was dirty, and overwrote the pending ARC buffer with its contents, and so those zeroes were written down. Whatever the bug is, this is the heart of it.</p> <p>Another interesting factor is that last twoskip record in the previous 128K ZFS record ends exactly on at the end of the record. This means that the map wouldn’t have been extended, as it was already right on the end of the block, covering the entire file. So the very first write to the new record would have been the one that allocated it, and the map extension immediately after would have been the first that covered that record, so the record would have only existed in any form in the ARC at that moment, not yet on disk.</p> <p>(This is interesting, because complex interactions at object boundaries are a traditionally a rich source of bugs).</p> <p>But the thing we must remember is that we have been running this code for nearly two years. We have written billions, maybe even trillions, of records in that time. If this was a simple case of a small handful of conditions occurring at the same time, we’d almost certainly have hit this already, and we know we haven’t. That almost certainly points to a locking bug; ZFS is working to keep the ARC and the page cache in sync, which means knowing their states in a given moment and choosing whether or not to copy data back and forth. It very much feels like there is a tiny, probably microsecond-long gap, where something is unlocked that shouldn’t be, and if you’re very very unlucky, a decision is made right in the gap based on faulty knowledge.</p> <p><strong>Making a test case</strong></p> <p>So Bron and I got to the point where we at least felt like we knew the moving parts that were likely involved, and we could start thinking about a test case. To reproduce a bug that only happens in a very specific set of conditions that you can control and one that you can’t, you just have to write a program that hits those conditions as often as possible, and then run it over and over in a short space of time, trying to hit the elusive condition that explodes the whole thing.</p> <p>So, Bron wrote <a href="https://github.com/cyrusimap/cyrus-imapd/tree/skipwork/contrib/skipwhack" target="_blank" rel="noopener">skipwhack</a>, which attempts to write to a file in the shape described above over and over, and then check back to see that it actually wrote stuff out properly.</p> <p>Now, at time of writing, I haven’t had chance to exercise this - all of my test systems have some deficiency (disks too slow, CPU too slow, not enough memory) that make it unable to generate enough load to trigger the bug. Of course, it could be that its also not an accurate enough reflection of the situation (maybe we missed a condition) to trigger it, but I think its far too soon to call that. I’ll be trying to find a suitable machine this week.</p> <p><strong>Maybe none of this matters</strong></p> <p>Around the time Bron emailed me a PR appeared that I think will probably fix this anyway: https://github.com/openzfs/zfs/pull/14498. It’s claiming to fix the previous mmap bug I mentioned upthread. I don’t really follow it well, but it seems like ZFS was expecting a page fault to either succeed or fail, but there’s a rare “in between” state that it wasn’t handling properly. This fix seems to make it handle that case.</p> <p>I suspect this is the whole story, but I’m not totally sure, I’m still going to be trying to get skipwhack to run and blow up, because then we can try this patch and see if it stops doing that. If it does, great, if not, I can write this up in more detail as there may be further things. We’ll see!</p> </blockquote><h3 id="postscript" tabindex="-1">Postscript</h3><p>We never saw another one of these corruptions, and we’re now on a newer ZFS which includes that fix, so we believe we’re all good for this. But in all these years, the only twoskip corruption we have seen has been an operating system level filesystem corruption!</p><p>Yes, we did blow away and recreate the filesystem which had given us the error as well. Our hardware really is super reliable these days, so we really don’t see this kind of issue very often.</p><p>Rob N is now working for <a href="https://klarasystems.com/" target="_blank" rel="noopener">Klara Systems</a>, on the OpenZFS filesystem. He writes more great material about it over on <a href="https://despairlabs.com/blog/" target="_blank" rel="noopener">his blog</a>. We’re all very happy to be using a filesystem that he’s working on improving!</p><h2 id="replacing-twoskip" tabindex="-1">Replacing twoskip</h2><p>Having said all this, the issues that we identified even back in 2016 still exist. Twoskip is a little high on random IO for writes. That’s OK for us. But we have single file databases with over a gigabyte; and the “checkpoint” command, where a file gets rewritten from scratch, can take many minutes. That’s painful, because it’s a stop the world lock.</p><p>So I’m working on a <a href="https://github.com/cyrusimap/cyrus-imapd/pull/5157" target="_blank" rel="noopener">new format</a>, <code>twom</code> - pronounced “tomb”. It’s basically what I described back in 2016:</p><ul> <li>Ancestor pointers so you can perform an <a href="https://en.wikipedia.org/wiki/Multiversion_concurrency_control" target="_blank" rel="noopener">MVCC</a> read without blocking writes</li> <li>a checkpoint operation which uses MVCC plus log replay, without ever needing long locks</li> <li>a faster hash algorithm (<a href="https://xxhash.com/" target="_blank" rel="noopener">xxHash</a> rather than CRC32)</li> <li>direct mmap reads and writes rather than mixing mmap with file IO</li> <li>blank slop space on the end of files to avoid needing so many syscalls to extend the file</li> <li>more direct file handle and mmap manipulation rather than using our library wrappers, so we can be smarter about duplicating key records.</li> </ul><p>That last one this is a bigger deal than you might imagine. We spend a lot of CPU in twoskip just duplicating data that could be only be lost in rare edge cases. By being more lazy about that, we can save CPU in the common case.</p><p>All this needs to be done without losing any of the existing high reliability that twoskip gives us. I’m very much hoping to have twom stable by early next year and start testing it in production!</p><h2 id="starvation" tabindex="-1">Starvation</h2><p>The thing that triggered me coming back to this problem was an archive user having a couple of million emails deleted. One week later, we run the <code>cyr_expire</code> tool to clean up the delete email, and this tool currently takes a lock and then does all the deletes. A million updates to a conversations database takes a long time, even just for deletes! There’s indexes to re-calculate and update, files to unlink.</p><p>It took a couple of hours, and while it did so, it blocked deliveries for that one user. Archiveusers for big customers get a lot of mail, this started using up all the processes and slowing email delivery for about 8% of our customers. Ouch. We paused delivery to that one user and got everyone else back working, but that’s annoying having to manually intervene, particularly outside working hours.</p><p>So we’re also working out way through the Cyrus code looking for things like this, and batching them up. In future, <code>cyr_expire</code> is going to do a bunch of messages, release all its locks, then come back and start again with another batch. That way, other tasks like delivering a new email get a chance to interleve.</p><h2 id="fastmail-really-cares-about-the-detail" tabindex="-1">Fastmail really cares about the detail</h2><p>Hopefully you can see from this, we’re right down in the weeds looking at exactly what’s going on, and building components with the kind of reliability that allows us to keep the service up, and keep email flowing, even as we run into nasty edge cases with our systems.</p><p>See you again in the next one.</p></content> 704 </entry><entry> 705 <title>Dec 5: MySQL InnoDB innodb_trx is cached</title> 706 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/mysql-innodb-trx-id/' /> 707 <id>https://www.fastmail.com/blog/mysql-innodb-trx-id/</id> 708 <updated>2024-12-05T00:00:01Z</updated><author> 709 <name>Rob Mueller</name> 710 </author><content xml:lang='en' type='html'><p>This is the fifth post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/meet-the-team-bek/">Dec 4: Meet the team — Bek</a>. The next post is <a href="/blog/twoskip-and-more/">Dec 6: Twoskip and more</a>.</p><p>This is a technical post about an aspect of MySQL InnoDB and transaction ids.</p><p>One of the features of <a href="https://datatracker.ietf.org/doc/html/rfc8620" target="_blank" rel="noopener">JMAP</a> is that it allows a client to <a href="https://datatracker.ietf.org/doc/html/rfc8620#section-5.2" target="_blank" rel="noopener">fetch changes</a> that have occurred on the server since the client last synced with the server. This is done with a <code>sinceState</code> string. Although no specific implementation is required, we’ve found that using a system based on “modification sequences” (aka MODSEQs) as originally described in the <a href="https://datatracker.ietf.org/doc/html/rfc4551" target="_blank" rel="noopener">IMAP CONDSTORE</a> extension works well.</p><p>I have been doing some work internally to allow us to calculate MODSEQs for all JMAP objects stored in our MySQL database. The basic idea was to setup triggers on any INSERT/UPDATE/DELETE actions to update the appropriate MODSEQ on the corresponding table row. The exact structures required for implementing the JMAP /changes method on all database tables is something for another blog post, but this post is more about an unexpected oddity I found while trying to do this.</p><p>Conceptually what I need is reasonably straightforward. There is a table to store all the current MODSEQs for each User owned table tuple.</p><pre><code>CREATE TABLE UserModSeqs ( 711 UserId INT NOT NULL, 712 TableName VARCHAR(255) NOT NULL, 713 CurrentModSeq BIGINT NOT NULL DEFAULT 0, 714 HighestPurgedModseq BIGINT NOT NULL DEFAULT 0, 715 PRIMARY KEY (UserId, TableName), 716 CONSTRAINT UserModSeqsFK FOREIGN KEY (UserId) 717 REFERENCES Users (UserId) ON DELETE CASCADE 718 ); 719 </code></pre><p>There is a generic “bump” function to increment the current MODSEQ for a particular User/table</p><pre><code>CREATE FUNCTION BumpUserModSeq (ModSeqOwnerValue INT, DataTableName VARCHAR(255)) 720 RETURNS BIGINT 721 BEGIN 722 DECLARE NewModSeq BIGINT; 723 724 INSERT INTO UserModSeqs 725 (UserId, TableName, CurrentModSeq, HighestPurgedModseq) 726 VALUES 727 (ModSeqOwnerValue, DataTableName, 1, 0) 728 ON DUPLICATE KEY UPDATE 729 CurrentModSeq = CurrentModSeq + 1; 730 731 SELECT CurrentModSeq 732 FROM UserModSeqs 733 WHERE UserId = ModSeqOwnerValue 734 AND TableName = DataTableName 735 INTO NewModSeq; 736 737 RETURN NewModSeq; 738 END 739 </code></pre><p>And then the actual trigger which looks something like:</p><pre><code>CREATE TRIGGER ${table}UpdateModSeq 740 BEFORE UPDATE ON $table 741 FOR EACH ROW 742 BEGIN 743 SET NEW.UpdatedModSeq = BumpUserModSeq(&quot;UserId&quot;, &quot;$table&quot;); 744 END 745 </code></pre><p>That you’d create for each <code>$table</code> that’s “owned” by a User (i.e. has foreign key UserId to the Users table).</p><p>One of the issues with this is that every single row updated on a table generates a new MODSEQ for each updated row.</p><p>I had an idea to make it so we only bump the modseq number once for each transaction rather than each row. Searching around you can find that the <code>information_schema.innodb_trx</code> table has a <code>trx_id</code> field, so a query like:</p><pre><code>SELECT trx_id 746 FROM information_schema.innodb_trx 747 WHERE trx_mysql_thread_id = connection_id() 748 </code></pre><p>Will get the transaction id of your current session, great. So we can create a function like:</p><pre><code>CREATE FUNCTION GetCurrentTrxId () 749 RETURNS BIGINT UNSIGNED 750 BEGIN 751 DECLARE CurrentTrxId BIGINT UNSIGNED; 752 753 SELECT trx_id 754 FROM information_schema.innodb_trx 755 WHERE trx_mysql_thread_id = connection_id() 756 INTO CurrentTrxId; 757 758 RETURN CurrentTrxId; 759 END 760 </code></pre><p>And then change the “bump” function to:</p><pre><code>CREATE FUNCTION BumpUserModSeq (ModSeqOwnerValue INT, DataTableName VARCHAR(255)) 761 RETURNS BIGINT 762 BEGIN 763 DECLARE NewModSeq BIGINT; 764 DECLARE CurrentTrxId BIGINT UNSIGNED DEFAULT GetCurrentTrxId(); 765 766 IF CurrentTrxId IS NULL OR 767 @LastBumpUserModSeqTrxId IS NULL OR 768 @LastBumpModSeqOwnerValue IS NULL OR 769 @LastBumpDataTableName IS NULL OR 770 @LastBumpUserModSeqTrxId &lt;&gt; CurrentTrxId OR 771 @LastBumpModSeqOwnerValue &lt;&gt; ModSeqOwnerValue OR 772 @LastBumpDataTableName &lt;&gt; DataTableName 773 THEN 774 775 INSERT INTO UserModSeqs 776 (UserId, TableName, CurrentModSeq, HighestPurgedModseq) 777 VALUES 778 (ModSeqOwnerValue, DataTableName, 1, 0) 779 ON DUPLICATE KEY UPDATE 780 CurrentModSeq = CurrentModSeq + 1; 781 782 # InnoDB may only create trx_id on first write (which we just did) 783 SET @LastBumpUserModSeqTrxId = IFNULL(CurrentTrxId, GetCurrentTrxId()); 784 SET @LastBumpModSeqOwnerValue = ModSeqOwnerValue; 785 SET @LastBumpDataTableName = DataTableName; 786 END IF; 787 788 SELECT CurrentModSeq 789 FROM UserModSeqs 790 WHERE UserId = ModSeqOwnerValue 791 AND TableName = DataTableName 792 INTO NewModSeq; 793 794 RETURN NewModSeq; 795 END 796 </code></pre><p>That way we bump the MODSEQ for the given object and user if it’s a new transaction, but use the existing value if we’re in a transaction where we already bumped the MODSEQ.</p><p>Now if you try this by hand a bit, it all seems to work great.</p><p>But, if you start trying to write tests for this, you start noticing it doesn’t always work as expected. Multiple updates in quick succession don’t bump the modseq correctly.</p><p>And so down the rabbit hole you go.</p><p>The first thing you learn is that InnoDB has an optimisation where it won’t generate a transaction id until you actually perform either a write statement or a <code>SELECT ... FOR UPDATE</code> (that’s the “InnoDB may only create trx_id on first write (which we just did)” comment in the above code).</p><pre><code>mysql (root@127.0.0.1) [fastmail]&gt; begin; select GetCurrentTrxId(); commit; 797 Query OK, 0 rows affected (0.00 sec) 798 799 +-------------------+ 800 | GetCurrentTrxId() | 801 +-------------------+ 802 | NULL | 803 +-------------------+ 804 1 row in set (0.00 sec) 805 </code></pre><pre><code>mysql (root@127.0.0.1) [fastmail]&gt; begin; select count(*) from Users where UserId=7 for update into @foo; select GetCurrentTrxId(); commit; 806 Query OK, 0 rows affected (0.00 sec) 807 808 Query OK, 1 row affected (0.00 sec) 809 810 +-------------------+ 811 | GetCurrentTrxId() | 812 +-------------------+ 813 | 4769402 | 814 +-------------------+ 815 1 row in set (0.00 sec) 816 </code></pre><p>Fine, but then you notice strange behavior like.</p><pre><code>mysql (root@127.0.0.1) [fastmail]&gt; begin; select GetCurrentTrxId(); select count(*) from Users where UserId=7 for update into @foo; select GetCurrentTrxId(); commit; 817 Query OK, 0 rows affected (0.00 sec) 818 819 +-------------------+ 820 | GetCurrentTrxId() | 821 +-------------------+ 822 | NULL | 823 +-------------------+ 824 1 row in set (0.00 sec) 825 826 Query OK, 1 row affected (0.00 sec) 827 828 +-------------------+ 829 | GetCurrentTrxId() | 830 +-------------------+ 831 | NULL | 832 +-------------------+ 833 1 row in set (0.01 sec) 834 835 Query OK, 0 rows affected (0.00 sec) 836 </code></pre><p>So even though a <code>SELECT ... FOR UPDATE</code> should create a transaction id, the subsequent call to <code>GetCurrentTrxId()</code> still didn’t return one.</p><p>But then if you try.</p><pre><code>mysql (root@127.0.0.1) [fastmail]&gt; begin; select GetCurrentTrxId(); select sleep(1) into @foo; select count(*) from Users where UserId=7 for update into @foo; select GetCurrentTrxId(); commit; 837 Query OK, 0 rows affected (0.00 sec) 838 839 +-------------------+ 840 | GetCurrentTrxId() | 841 +-------------------+ 842 | NULL | 843 +-------------------+ 844 1 row in set (0.00 sec) 845 846 Query OK, 1 row affected (1.00 sec) 847 848 Query OK, 1 row affected (0.00 sec) 849 850 +-------------------+ 851 | GetCurrentTrxId() | 852 +-------------------+ 853 | 4770035 | 854 +-------------------+ 855 1 row in set (0.00 sec) 856 857 Query OK, 0 rows affected (0.00 sec) 858 </code></pre><p>So putting a sleep between them, suddenly you do get a new transaction id.</p><p>Digging into MySQL code you end up finding</p><p><a href="https://github.com/mysql/mysql-server/blob/61a3a1d8ef15512396b4c2af46e922a19bf2b174/storage/innobase/handler/i_s.cc#L751" target="_blank" rel="noopener">i_s.cc#L751</a></p><pre><code>/** Common function to fill any of the dynamic tables: 859 INFORMATION_SCHEMA.innodb_trx 860 @return 0 on success */ 861 static int trx_i_s_common_fill_table( 862 ... 863 /* update the cache */ 864 trx_i_s_cache_start_write(cache); 865 trx_i_s_possibly_fetch_data_into_cache(cache); 866 </code></pre><p><a href="https://github.com/mysql/mysql-server/blob/61a3a1d8ef15512396b4c2af46e922a19bf2b174/storage/innobase/trx/trx0i_s.cc#L817" target="_blank" rel="noopener">trx0i_s.cc#L817</a></p><pre><code>int trx_i_s_possibly_fetch_data_into_cache( 867 trx_i_s_cache_t *cache) /*!&lt; in/out: cache */ 868 { 869 if (!can_cache_be_updated(cache)) { 870 return (1); 871 } 872 </code></pre><p><a href="https://github.com/mysql/mysql-server/blob/61a3a1d8ef15512396b4c2af46e922a19bf2b174/storage/innobase/trx/trx0i_s.cc#L681" target="_blank" rel="noopener">trx0i_s.cc#L681</a></p><pre><code>static bool can_cache_be_updated(trx_i_s_cache_t *cache) /*!&lt; in: cache */ 873 { 874 ... 875 /** The minimum time that a cache must not be updated after it has been 876 read for the last time. We use this technique to ensure that SELECTs which 877 join several INFORMATION SCHEMA tables read the same version of the cache. */ 878 constexpr std::chrono::milliseconds cache_min_idle_time{100}; 879 880 return std::chrono::steady_clock::now() - cache-&gt;last_read.load() &gt; 881 cache_min_idle_time; 882 </code></pre><p>So the <code>information_schema.innodb_trx</code> table is cached internally for 100ms, so trying to use it to fetch the transaction for the current connection is not guaranteed to be up to date. Ouch! This does not appear to be documented anywhere obvious that I could find.</p><p>Unfortunately I can’t see any way to <em>reliably</em> find out if you’re in a transaction or way to identify the current transaction id.</p><p>In the end, I had to resort to something at the application level. Internally our DB library encourages you to use a guard object to start and end transactions.</p><pre><code>my $committer = $dbh-&gt;begin_work_auto(); 883 $dbh-&gt;do(...); 884 $committer-&gt;commit(); 885 </code></pre><p>The <code>-&gt;begin_work_auto()</code> call starts a transaction and returns a guard object. If the <code>$committer</code> guard object is destroyed before the <code>-&gt;commit()</code> call, then it will automatically rollback the transaction in progress.</p><p>I added a small bit of code in <code>begin_work_auto()</code> to do <code>$self-&gt;do('SET @CurrentTrxId=?', {}, ++$count);</code>, and then set it back to NULL when we commit/rollback. That simulates a guaranteed changing transaction id within each MySQL session as long as you use the <code>begin_work_auto()</code> guard method which most code does when using transactions. Anything that doesn’t use <code>begin_work_auto()</code> leaves <code>@CurrentTrxId</code> as NULL.</p><p>The <code>BumpUserModSeq</code> function used in the trigger checks the <code>@CurrentTrxId</code> session variable to see if it’s NULL or has been incremented from it’s last value to decide whether to bump the MODSEQ or not.</p><p>So transactions that use <code>begin_work_auto()</code> get optimised MODSEQ bumping (a single update to the modseq table for all rows added/updated/deleted for a user within the transaction), and everything else falls back to still working correctly but with a bit more overhead (an update to the modseq table for every single row added/updated/deleted).</p></content> 886 </entry><entry> 887 <title>Dec 4: Meet the team — Bek</title> 888 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/meet-the-team-bek/' /> 889 <id>https://www.fastmail.com/blog/meet-the-team-bek/</id> 890 <updated>2024-12-04T00:00:01Z</updated><author> 891 <name>The Fastmail Team</name> 892 </author><content xml:lang='en' type='html'><p>This is the fourth post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/moving-house-new-datacentre/">Dec 3: On moving house — bringing a new data centre online</a>. The next post is <a href="/blog/mysql-innodb-trx-id/">Dec 5: MySQL InnoDB innodb_trx is cached</a>.</p><p>Meet Bek, our Head of People and Culture based in Melbourne.</p><p><strong>Name:</strong> Bek Fraser</p><p><strong>Role:</strong> Head of People and Culture</p><p><strong>What do you work on?</strong></p><p>Anything people related, including all our policies, processes, compliance. All the really boring stuff. Also, how to keep our team members happy.</p><p><strong>How long working at Fastmail, how did you get involved?</strong></p><p>10 months. I was looking for a role in an IT company. I heard of Fastmail many years ago when I first started my IT degree, and Fastmail was leading the email industry.</p><p>I’m Fastmail’s first people and culture hire. The exciting part was starting from a blank slate — it’s been a good opportunity to look at how to align our teams, which are global with many remote team members.</p><p><strong>What’s a project you have worked on that you’re particularly proud of?</strong></p><p>Implementation of a skill and career framework which has created equality for all our team members globally, and is allowing us to look at what skills we need for the future.</p><p><strong>Who is somebody who inspires you?</strong></p><p>Early on in life it was Mother Theresa, because she just wanted to help everybody — but she didn’t hesitate to break the mould. Something I learned being a female in IT, you never can hesitate to break the mould.</p><p><strong>What are your favourite Fastmail features?</strong></p><p>Memos! It means I can keep information about something without having to go out to a separate tool. It took me a while to get used to conversation mode, but I’ve found it useful to help my husband manage his time better — I’ve taught him how to use conversations and colour coding to keep track of things.</p><p><strong>Other than Fastmail, what’s your favourite or most used piece of technology?</strong></p><p>My robot that mops and vacuums at the same time! It notices when it’s on carpet and lifts the mop. It saves me tons of time and I can vacuum the floor every day.</p><p><strong>What are you listening to these days?</strong></p><p>John Farhnam — I did the backing for Age of Reason! You can see tiny little me in the music video. I also enjoy Guy Sebastian, Adele, that type of music. I also love musicals. I’m old school, I love music that wasn’t created by computer. My grandfather was a concert pianist.</p><p><strong>What do you like to do outside of work?</strong></p><p>I like to challenge myself. I built outside stairs last week, now I need to repaint the whole deck to match them. I love working in the garden. I make resin chopping boards and sell them at the market.</p><p><strong>What’s your favourite animal?</strong></p><p>My puppy dog. I love dogs, especially my puppy Baxter.</p><p><strong>Any Fastmail staff you want to brag on?</strong></p><p>All of them. We’ve achieved a lot in the last 12 months, and everyone really lives the Fastmail culture.</p><p><strong>What do you like best about working at Fastmail?</strong></p><p>The people and their commitment. It’s been great helping everyone find the type of growth that works best for them, and creating a culture at Fastmail that truly aligns with what we believe in.</p></content> 893 </entry><entry> 894 <title>Dec 3: On moving house — bringing a new data centre online</title> 895 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/moving-house-new-datacentre/' /> 896 <id>https://www.fastmail.com/blog/moving-house-new-datacentre/</id> 897 <updated>2024-12-03T00:00:01Z</updated><author> 898 <name>Graeme Lee</name> 899 </author><content xml:lang='en' type='html'><p>This is the third post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/throwback-security-confidentiality-integrity-availability/">Dec 2: Thowback: security — confidentiality, integrity, and availability</a>. The next post is <a href="/blog/meet-the-team-bek/">Dec 4: Meet the team — Bek</a>.</p><p>Who hasn’t moved house at some point in their lives? Whether we choose to rent or buy, board, or backpack, each relocation in premises has its own challenges which have to be considered.</p><p>You might be striking out, leaving the nest and finding your own dwelling. Or you may have been comfortably in the one location for many years, and now it’s time to move on. Maybe you live in a motor home? Then relocating is probably something you <em>just do</em>. It’s different for everyone, and the decisions and circumstances are unique to each individual.</p><p>Well, for reasons, Fastmail decided it was time to relocate our DC’s. And it felt a lot like moving house. How? Well, here are some cross-overs.</p><ul> <li>How much space do we need? Usually the first question asked, I suppose. Do we have enough already? Are we wasting space? What if we need more? How do we expand? If the family has moved out, we can downsize. Or there might be another project on the way, and we need another room.</li> <li>What’s the neighbourhood like? When house-hunting, we will consider how we might fit in with the community we are integrating into. We will look at services, schools, shopping facilities. Whatever we may prioritise, it will be unique to each one’s circumstances. When we look at data centres, we might consider their physical location. What about connectivity? Do we have specific requirements for any services? What are the facilities like when connecting, upgrading, or disconnecting?</li> <li>What’s the access like? Is it miles from anywhere, or is it in the CBD? Are we prepared to commute, or do we want quick and easy access? Does it have a garage or is it street parking only? The same questions can be asked when choosing a new DC. We might appreciate the opportunity to get out and visit. Or do we want something close that has convenient parking.</li> <li>Is the furniture we have going to work in our new abode? Maybe it’s all we need, and we’ll “make it fit”. Do we have the budget to refurbish? What can we throw away? What do we keep? We could go on-and-on here. But it’s important to accept that what may work for one person may not necessarily be the right fit for the next. The same goes for moving to a new data centre. You might need new hardware, so it’s time to find a new space and start fresh. Whatever. What you choose to bring and what you choose to leave behind is always your choice alone.</li> </ul><h2 id="timing-is-everything" tabindex="-1">Timing is everything</h2><p>How long will the move take? Maybe you’re moving across the continent, so it’s going to take a few days travel by road. Or it’s just around the corner, so logistically, the move looks like it will be fairly quick. But now, you’re moving data centres. The lights need to stay on! We now have a conundrum. Your servers might be fine with being turned off. But not your service!</p><p>How do you go about moving a live production system from one data centre to another with zero downtime as the goal? This is where your existing data integrity, resilience and backup strategies come in to play.</p><p>At Fastmail, data integrity is of utmost importance. It was time to take our integrity policy to task! We were about to stretch our network’s capabilities to the limit, break some old rules and make new ones, and find some new corner cases that we had not considered.</p><p>And you can’t just do “1 thing”, and move on to the next. A lot of things need to be thought out and prepared for before you can turn up at the doorstep and ask for the keys!</p><p>We did the tours. We shopped around. And we found a data centre that we felt good about. It was accessible, and had the right demographic of network carriers in it for our needs. Our floorplan was drawn up, cables were run, and the plumbing was good to go.</p><h2 id="let-s-do-this" tabindex="-1">Let’s do this!</h2><p>New switching and routing hardware was delivered and installed. We needed some way to bridge the two data centres together, so we used <a href="https://www.megaport.com/" target="_blank" rel="noopener">Megaport</a> to provide a private circuit between our sites so that we could transparently continue to provide services from our existing network connections without disrupting normal operations.</p><p>We decided to divide and conquer our data. We already have data replicated between our primary and backup DC’s. It made sense to employ a similar technique. We transferred the operational load onto the servers that were staying up, offlined our standby systems, put them on a truck, and shipped. Things were ok. But as time passed, we discovered that the load now on our remaing hosts was taxing them to their limits! Fortunately, the time-in-transit was only a few hours, and we were able to keep things running.</p><p>Once our standby hosts were in place, we began to re-sync them with our master systems. This was fairly straightforward. But we quickly discovered that we had not provisioned enough bandwidth between our DC’s to cater for the high intra-network load. Extra bandwidth on our DC cross-connect was provisioned, and things synced up nicely.</p><p>Because we experienced a higher than expected load on our existing servers, we had to re-think our redundancy strategy. We had already moved 2/3 of our compute, with 1/3 remaining (excluding our backup DC), and we felt that it wasn’t desirable to rely on 1/3 in the event of an emergency, so a rethink was in order. Our server pool was reconfigured to provide more compute and more resilience to our primary DC.</p><p>Finally, we brought the remainder of our servers across. This was possibly the most straight-forward part of the move. Our new DC was promoted to master, and the servers were offlined, shipped, and installed.</p><p>Happy days!</p><h2 id="but-wait-there-s-more" tabindex="-1">But wait, there’s more!</h2><p>A lot more! TL;DR more (At this point, does ‘TL’ mean ‘Too Late!’?) It wasn’t plain sailing. We had some hurdles that had to be swiftly cleared so that we could continue with our migration. Cables didn’t arive. We ran out of certain types of SFP modules. We upgraded every NIC on every blade. And we managed to migrate our backup DC right on the heels of our primary!</p><p>I’m sure we could write another post or two about these things. But we are still happy with the outcome. Some of our tooling got a fresh look at under stress, and we were able to improve even further our backup and migration tools for both present and future use.</p><p>The whole global team stepped up for this, and as a result, we were successful. Our network has been revitalised for the future, and we are in a great position to grow and improve our backend to provide the excellent service our customers expect and deserve.</p><p>No emails were harmed or lost in the migration of our data centres!</p><hr><blockquote> <p>Postscript by Bron: the whole team did amazing work here. While we didn’t lose any email, we did have some customer visible downtime unfortunately. This all happened in just a few weeks - shipping equipment from Seattle to Philly over a week first, then the crazy day where we were running fully split as we shipped half of New Jersey to Philly, then finally I was off at a conference while the remaining crew handled the move of the last parts from New Jersey to our new secondary location.</p> <p>We identfied three major learning points from this move:</p> <ol> <li>We hadn’t done enough testing of the states we were going to be in during the move. We believed things would work that turned out to not be able to handle the complete real-world load during the US daytime hourly spikes.</li> <li>We didn’t actually have the capacity! We also had a Cyrus replication repair inefficiency which, along with a very inopportune IMAP server crash due to the extra load of running overloaded - caused days of scrambling to get things back into the right state. We didn’t lose any data, but we were running with lower redundancy than I wanted for longer than I wanted! We have ordered a bunch of new servers which were shipped just last week and will be up and running by the end of the year.</li> <li>There was no single coordinator driving the entire process. We still need to be better at designating a single responsible person. With such a senior and self-directing team it wasn’t a disaster, but we could have been even more efficient with a bit more up-front planning.</li> </ol> <p>There were also some network hiccups that happened in the weeks immediately following; but we’re now in a really stable configuration, with more options for routing traffic than we’ve ever had before, so we can work around faults faster in future.</p> </blockquote></content> 900 </entry><entry> 901 <title>Dec 2: Throwback: security — confidentiality, integrity, and availability</title> 902 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/throwback-security-confidentiality-integrity-availability/' /> 903 <id>https://www.fastmail.com/blog/throwback-security-confidentiality-integrity-availability/</id> 904 <updated>2024-12-02T00:00:01Z</updated><author> 905 <name>Bron Gondwana</name> 906 </author><content xml:lang='en' type='html'><p>This is the second post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The previous post was <a href="/blog/mission-statement/">Dec 1: Mission statement</a>. The next post is <a href="/blog/moving-house-new-datacentre/">Dec 3: On moving house — bringing a new data centre online</a>.</p><p>Throwback time! <a href="/blog/security-confidentiality-integrity-and-availability/">This was the post</a> which inspired our first ever Advent series. All I’ve changed is putting an Oxford comma in the title.</p><p>Honestly, very little has changed since then. Even the Wikipedia link is still valid. The James Mickens paper has disappeared, Microsoft’s legendary commitment to backwards compatibility clearly doesn’t extend that far, so if you’re interested you’ll have to <a href="https://scholar.harvard.edu/files/mickens/files/thisworldofours.pdf" target="_blank" rel="noopener">get the file from Harvard</a> (pdf). I strongly recommend reading anything he has written.</p><h2 id="integrity" tabindex="-1">Integrity</h2><p>We did an <a href="/blog/security-integrity/">explicit post on integrity in 2014</a>.</p><p>I’m going to write a whole separate post in this series about how the data integrity checks we have added over the years have really held up. We haven’t had a major data loss in those 10 years, through many cases of equipment failure and the occasional human error. Designing with data resilience as a key goal has paid off.</p><h2 id="availability" tabindex="-1">Availability</h2><p>We also wrote an <a href="/blog/security-availability/">explicit post on availability in 2014</a>.</p><p>Sadly we haven’t been with NYI for a while. Miss those guys. We moved to New Jersey, then they sold that data centre to a new provider, and we were there for a while. Our current data centres are in Philadelphia and St Louis.</p><p>We did a major data centre move earlier this year. We will write in this series about the experience, some lessons learned about our preparedness, and how the move has improved our resilience against some of the risks out there. Unfortunately, it did lead to some higher levels of downtime during and immediately after the move as things settled.</p><h2 id="confidentiality" tabindex="-1">Confidentiality</h2><p>Finally the first one everybody thinks about! <a href="blog/security-confidentiality">We wrote about confidentiality in 2014</a> as well.</p><p>The reasoning here is still the same, though a few things have changed. Our data centre structure is a bit different, we handle the networking in-house now and use a single switching infrastructure with VLANs rather than airgapped networks. The legal framework has changed too, Australia now has a Cloud Act agreement with the USA, so we receive requests directly from the USA for investigations of serious crimes by non-Australians.</p><p>By far the biggest risk to confidentiality that didn’t exist 10 years ago to the same extent is AI training! It seems that the temptation is for services that are “freemium” to <a href="https://www.linkedin.com/help/linkedin/answer/a5538339" target="_blank" rel="noopener">use your</a> <a href="https://slack.com/intl/en-au/trust/data-management/privacy-principles" target="_blank" rel="noopener">data for</a> <a href="https://www.cnet.com/tech/services-and-software/how-to-opt-out-of-instagram-and-facebook-using-your-posts-for-ai/" target="_blank" rel="noopener">training their</a> <a href="https://support.google.com/mail/answer/10079371" target="_blank" rel="noopener">AI models</a>, either without choice or by forcing you to find a <a href="https://www.goodreads.com/quotes/40705-but-the-plans-were-on-display-on-display-i-eventually" target="_blank" rel="noopener">well hidden</a> way to opt out.</p><p>We aren’t doing any AI, and if we did it would be very carefully and with the only goal being to help our users get better insight into their email. We would use per-user training models, in the same way that we currently create per-user search indexes, ensuring that data is segregated and can’t leak across.</p><p>The great thing about having a paid product is that we don’t have split loyalties. Fastmail has been profitable every one of the past 10 years, with no outside investment. We expect to remain so into the future without having to “diversify” into shady shit. We sleep well knowing we run an ethical business, and we’re very grateful to the people who trust us with their email and pay us to keep providing them with the service.</p><hr><p>In conclusion, Fastmail has exactly the same attitude to security that we had 10 years ago. It’s important. It’s not a marketing dot-point, it’s table stakes — you have to be secure if you want to be trusted with other people’s precious emails.</p><p>Other than the areas where we’re the ones creating new and better standards, we are cautious about adopting new technologies and the latest fads. This has served us very well over the years. Fastmail cares about real security, actionable changes that make things better. We don’t do security theatre. When we do something in the name of security, it’s because it makes a meaningful difference to your risk profile.</p><p>See you again tomorrow.</p></content> 907 </entry><entry> 908 <title>Dec 1: Mission statement</title> 909 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/mission-statement/' /> 910 <id>https://www.fastmail.com/blog/mission-statement/</id> 911 <updated>2024-12-01T00:00:01Z</updated><author> 912 <name>Bron Gondwana</name> 913 </author><content xml:lang='en' type='html'><p>This is the first post in the <a href="/blog/fastmail-advent-2024/">Fastmail Advent 2024</a> series. The next post is <a href="/blog/throwback-security-confidentiality-integrity-availability/">Dec 2: Thowback: security — confidentiality, integrity, and availability</a>.</p><p>Mission statements have a bad reputation. For good reason. They’re generally somewhere between aspirational goal and corporate bullshit: don’t be evil, think different, enhance shareholder value…</p><p>So anyway. We decided to do one:</p><blockquote> <p>Make email better</p> </blockquote><p>The very first page in our Notion workspace starts with that phrase, and then explains it:</p><blockquote> <p>Fastmail is a small company making a big difference.</p> <p>We make email better for our customers by providing the email service that people are proud to pay for. And we make email better for the world by leading standards, open source, and advocacy work.</p> </blockquote><p>People who use our service know that the first item is true. Our customers are very loyal, for good reason. Sure, you can get email for free, if free is your only consideration, but our interface is smooth and fast, we have real humans staffing our support desk, and they sit RIGHT next to us. I am less than 2 metres away from the closest support agent right now.</p><p>Fastmail has its mission statement, and I have my own personal mission statement as well: “keep email open”. We had fully ⅓ of our engineering staff at the <a href="https://www.ietf.org/meeting/121/" target="_blank" rel="noopener">IETF meeting in Dublin</a> a few weeks ago. We’re heavily involved in <a href="https://www.ietf.org/about/introduction/" target="_blank" rel="noopener">developing the standards</a> that will keep email as the number one social network in the world, and we’re widely respected in the email industry.</p><p>I firmly believe that email is the <a href="https://www.fastmail.com/blog/email-is-your-electronic-memory/" target="_blank" rel="noopener">electronic memory</a> that we all need in a world where online content is changed frequently. I routinely find emails from 10 or even 20 years ago to remind me what happened then, and what I thought about it.</p><p>So there it is — Fastmail’s mission statement. We strive to live up to it every day. I judge myself as CEO, and prioritise our work, around meeting it. Creating a great product, and leading the world in making email better for everyone. I hope it comes through in every interaction everyone has with us, and you — the reader — agree with me that it’s how you see Fastmail as well.</p><p>See you tomorrow for the next post!</p></content> 914 </entry><entry> 915 <title>Fastmail Advent 2024</title> 916 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/fastmail-advent-2024/' /> 917 <id>https://www.fastmail.com/blog/fastmail-advent-2024/</id> 918 <updated>2024-12-01T00:00:00Z</updated><author> 919 <name>Bron Gondwana</name> 920 </author><content xml:lang='en' type='html'><p>Hello everyone, and welcome to Fastmail’s Advent blog post series for 2024!</p><p>After an experiment with <a href="/blog/fastmail-advent-2023-25-days-of-better-email/">Mastodon posts last year</a>, we’re back to our own platform, where we can do longer posts with multiple links again.</p><p>This is a special year for us: 10 years since our <a href="/blog/fastmail-advent-2014/">very first Advent blog post series</a>, and 25 years since the founding of Fastmail. More importantly, it’s my own 20 year anniversary; I started with FastMail (as we spelled it back then) as a sysadmin/programmer in October 2004!</p><p>In this series, you’ll see some throwbacks to posts we made 10 years ago, discussing what’s changed and what’s still the same. You’ll meet some of our newer staff members. You’ll see some of the documents that describe what we do and, more importantly, why we do it. What makes Fastmail special and why after all this time we’re still doing email and still passionate about making email better for everyone.</p><p>So pop our RSS feed into your reader, follow us on the socials, or just come back and visit our blog every day between your coffee and your wordle — or whatever guilty pleasure you prefer. Let’s get started:</p><ul> <li><a href="/blog/mission-statement/">Dec 1: Mission statement</a></li> <li><a href="/blog/throwback-security-confidentiality-integrity-availability/">Dec 2: Throwback: security — confidentiality, integrity, and availability</a></li> <li><a href="/blog/moving-house-new-datacentre/">Dec 3: On moving house — bringing a new data centre online</a></li> <li><a href="/blog/meet-the-team-bek/">Dec 4: Meet the team — Bek</a></li> <li><a href="/blog/mysql-innodb-trx-id/">Dec 5: MySQL InnoDB innodb_trx is cached</a></li> <li><a href="/blog/twoskip-and-more/">Dec 6: Twoskip and more</a></li> <li><a href="/blog/revision-of-core-email-specifications/">Dec 7: Revision of the core email specifications</a></li> <li><a href="/blog/principles/">Dec 8: Guiding principles</a></li> <li><a href="/blog/building-a-blog/">Dec 9: Building a blog</a></li> <li><a href="/blog/sunsetting-pobox/">Dec 10: Sunsetting Pobox</a></li> <li><a href="/blog/meet-the-team-marc/">Dec 11: Meet the team—Marc</a></li> <li><a href="/blog/following-the-sun/">Dec 12: Following the Sun</a></li> <li><a href="/blog/moving-fastmail-dns-to-knot/">Dec 13: It’s knot DNS. There’s no way it’s DNS. It is DNS!</a></li> <li><a href="/blog/on-call-systems/">Dec 14: On-call systems</a></li> <li><a href="/blog/platform-team-working-agreement/">Dec 15: Platform Team working agreement</a></li> <li><a href="/blog/offline-in-beta/">Dec 16: Offline support now in public beta</a></li> <li><a href="/blog/offline-architecture/">Dec 17: Building offline: general architecture</a></li> <li><a href="/blog/offline-sync/">Dec 18: Building offline: syncing changes back to the server</a></li> <li><a href="/blog/offline-mail-storage/">Dec 19: Building offline: mail storage</a></li> <li><a href="/blog/how-fastmail-uses-fastmail/">Dec 20: How Fastmail uses Fastmail!</a></li> <li><a href="/blog/fastmail-in-a-box/">Dec 21: Fastmail in a box</a></li> <li><a href="/blog/why-we-use-our-own-hardware/">Dec 22: Why we use our own hardware at Fastmail</a></li> <li><a href="/blog/ten-years-of-jmap/">Dec 23: Ten years of JMAP</a></li> <li><a href="/blog/twenty-five-years-of-fastmail/">Dec 24: Twenty five years of Fastmail</a></li> </ul></content> 921 </entry><entry> 922 <title>Introducing memos: stick private notes on your email</title> 923 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/introducing-memos/' /> 924 <id>https://www.fastmail.com/blog/introducing-memos/</id> 925 <updated>2024-08-29T00:00:00Z</updated><author> 926 <name>Neil Jenkins</name> 927 </author><content xml:lang='en' type='html'><p>Today we’re launching a new feature at Fastmail: memos.</p><p>Add a memo to an email to jot down private notes — remind yourself of tasks, record when you paid a bill, capture notes from a side conversation, or anything else!</p><p>Your memo will stay stuck to the top of the email, so you won’t forget. Want to change it? Just click and type, it saves automatically. It’s super simple, but surprisingly powerful.</p><p>Memos are private, only visible to you. They are not sent to anyone else in the conversation.</p><p>When you add a memo it will show up in the message list as well, so you can see at a glance the memos you’ve added to conversations in a label or folder. Here’s a quick look at it in action:</p><div class="mt-xl"><video class="aspect-video relative rounded-lg shadow-card z-50" playsinline muted controls width="2048" height="1152"> <source src="/assets/video/tour/Memos.webm" type="video/webm; codecs=vp9,vorbis"> <source src="/assets/video/tour/Memos.mp4" type="video/mp4"> </video></div><p>We’ve integrated memos deeply into our search as well. Text in your memos will be matched against words you search for, just like with your messages. You can search for text specifically inside memos with the operator <code>memo:&lt;text&gt;</code>. Or, want to just find all the messages with a memo? Search for <code>has:memo</code>.</p><p>If you also use another email app to access your Fastmail account, such as Apple Mail or Thunderbird, you’ll still have access to your memos. You’ll find them in the Memos folder, as a reply to the message your memo is attached to.</p><p>Memos are available in Fastmail now. Please note, if you have turned conversation grouping off, you’ll need to <a href="https://www.fastmail.help/hc/en-us/articles/1500001969861-Conversations#turnoff" target="_blank" rel="noopener">turn it back on</a> to use memos. In multi-user accounts, memos cannot be added to messages that are shared with you by another user; you can only add them to your own messages.</p><p>Not a Fastmail user yet? <a href="/features/">Discover all our powerful features</a> to make you more productive, or <a href="https://app.fastmail.com/signup/" target="_blank" rel="noopener">start your 30-day free trial today</a>.</p></content> 928 </entry><entry> 929 <title>Introducing passkey support to Fastmail</title> 930 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/introducing-passkeys/' /> 931 <id>https://www.fastmail.com/blog/introducing-passkeys/</id> 932 <updated>2024-08-14T01:30:00Z</updated><author> 933 <name>Neil Jenkins</name> 934 </author><content xml:lang='en' type='html'><p>From today, we’re pleased to announce you can create passkeys for your Fastmail account, giving you a faster, more secure way to log in. Open the <a href="https://app.fastmail.com/settings/security" target="_blank" rel="noopener">Privacy &amp; Security</a> settings in your account to get started, or read on to learn more about passkeys.</p><h2 id="the-problem-with-passwords" tabindex="-1">The problem with passwords</h2><p>To understand why passkeys are useful, we should start by looking at the problems with passwords that we want to solve.</p><p>A password is a shared secret between you and the website you’re logging into. You tell the website who you are (your username/email address) and the secret only you know, the website checks it matches the secret you gave them before, and if it’s a match then you’re logged in.</p><p>So far, so good. The trouble is, people are terrible at using passwords.</p><ul> <li> <p><strong>We’re not good at coming up with hard to guess passwords</strong>. In security, we have the concept of entropy, which is a mathematical way of calculating how predictable something is. The higher the entropy, the less predictable, and the stronger the security. This is important to stop people from being able to guess your password. Unfortunately, when most people try to come up with a password, they choose something with low entropy, which makes it susceptible to “brute force” attacks, where a computer can try billions of passwords very quickly until it finds which one is yours.</p> </li> <li> <p><strong>We’re not good at remembering lots of different passwords</strong>. As we saw above, a truly secure password is something that’s unpredictable. And if it’s unpredictable, it’s probably hard to memorise. Memorising one strong password is probably doable, but using the same password at every site is a terrible idea from a security point of view: all it takes is one website to store its passwords insecurely and a hacker could now access all your accounts, anywhere on the internet!</p> </li> <li> <p><strong>We’re not good at only giving our password to the right website</strong>. You may have a phenomenal memory, and maybe <a href="https://diceware.dmuth.org/" target="_blank" rel="noopener">roll dice to create a unique password with high entropy</a> for every account you create. Unfortunately, this doesn’t help if you click a link to a phishing website and hand it straight over to an attacker. This is the biggest cause of stolen Fastmail accounts by a long way.</p> </li> </ul><p>Now, you may be thinking there’s already an answer to the above: password managers. And that’s absolutely right. If there’s one thing you can do to improve your security on the web, it’s use a password manager. A password manager:</p><ul> <li>Creates secure, high-entropy passwords for you.</li> <li>Remembers them, so you don’t have to.</li> <li>Will only automatically fill it in on the website you created the password. (But as autofill isn’t 100% reliable, users can be tricked into manually copying the password out of the manager and into a phishing site).</li> </ul><p>That mostly solves our problems! But if we’re using a password manager anyway, we can make it even more secure by storing passkeys in it instead of passwords. All modern password managers also support passkeys (there’s a built-in one on every device these days, or we recommend <a href="https://1password.com" target="_blank" rel="noopener">1Password</a> for a good cross-platform password manager).</p><h2 id="what-are-passkeys-and-why-are-they-better-than-passwords" tabindex="-1">What are passkeys, and why are they better than passwords?</h2><p>Instead of a password - which is a shared secret with the website - a passkey a is a super-secure cryptographic key. It uses public key cryptography, which means your password manager stores a key that can prove it’s you, while the website gets a different key that can only <em>verify</em> this assertion. This has a number of advantages over passwords:</p><ul> <li><strong>It’s replay resistant</strong>. A password is the same every time, so if an attacker can observe it being sent to the website, they can use it themselves. With passkeys, you sign a different random “challenge” from the website every time, so even if an attacker can intercept your network traffic, they can’t steal your passkey and log in as you.</li> <li><strong>It’s database-leak resistant</strong>. If a website’s password database gets hacked, there’s a risk the attacker could get your password. (Good websites will have <a href="https://cheatsheetseries.owasp.org/cheatsheets/Password_Storage_Cheat_Sheet.html" target="_blank" rel="noopener">hashed the passwords</a> to slow the attackers down, but it’s still possible.) With passkeys, you have both a private key (that can create a signature to prove it’s you) and a public key (that can only verify the signature is real, but not create it). The website only ever has the public key, so even if this were stolen it couldn’t be used to access your account.</li> <li><strong>It’s phishing proof</strong>. Passkeys are strongly tied to the website they were created for. If you click a phishing link and end up on a malicious website, your passkey simply won’t appear, so you can’t get phished.</li> <li><strong>It’s quicker and easier</strong>. Browsers can provide secure, privacy-preserving APIs for integrating passkeys on a website to make logging in as easy as a single click.</li> </ul><p>Your password manager handles all of this for you. When you want to log in to a website or app, you just use your fingerprint, face, or PIN to verify it’s you (this happens entirely locally - your biometrics never leave your device). The website then receives a digital signature proving you have the passkey, and you’ll be logged in.</p><h2 id="how-do-i-log-in-on-a-new-device" tabindex="-1">How do I log in on a new device?</h2><p>Your passkey will be stored in your password manager. This will often automatically sync between all your devices. For example, if you use multiple Apple devices, the iCloud keychain will securely sync your passkeys to all of them.</p><p>If you don’t have the same password manager available on all your devices, you can also:</p><ul> <li>Use the passkey on your phone to log in on another device.</li> <li>Create another passkey on the other device to make it easier to log in there going forwards.</li> </ul><p>You can create as many passkeys for your account as you need; any of them can be used to log in. You can review and revoke the passkeys created for your account at any time from the <a href="https://app.fastmail.com/settings/security" target="_blank" rel="noopener">Privacy &amp; Security</a> settings.</p><h2 id="how-does-this-interact-with-my-existing-password-what-if-i-have-two-step-verification" tabindex="-1">How does this interact with my existing password? What if I have two-step verification?</h2><p>Passkeys are an additional way to log in, not currently a replacement for passwords. If you create a passkey, you will still be able to log in with your password, just as you did before. If you have two-step verification, this will still be required when you use your password. Two-step verification is not required when you use your passkey, as this already has two factors:</p><ul> <li>Something you have (the passkey on your device).</li> <li>Something you are or something you know (the touch/fingerprint or PIN your device requires to use it).</li> </ul><h2 id="the-state-of-passkeys-on-the-web-today" tabindex="-1">The state of passkeys on the web today</h2><p>We’re a big believer in passkeys, and hope they are the future for authentication. Phishing is a blight on the internet, and this is our best hope of eliminating it forever. We’ve actually had passkeys implemented for over a year at Fastmail, but we’ve been waiting for the ecosystem to mature before releasing it. While it’s come along way, there are still some rough edges.</p><p>The biggest issue is the integration of third-party password managers in browsers is still not as polished as it should be. Browser vendors are working on an API for passkey integration but it’s not widely supported yet, so instead the password managers are injecting JavaScript into the page to overwrite the native <a href="https://developer.mozilla.org/en-US/docs/Web/API/Navigator/credentials" target="_blank" rel="noopener"><code>navigator.credentials</code></a> object. The result is more fragile - for example, we ran into a bug with 1Password in Firefox, where the <a href="https://developer.mozilla.org/en-US/docs/Web/API/AuthenticatorAttestationResponse/getAuthenticatorData" target="_blank" rel="noopener">getAuthenticatorData</a> method returned an object that throws an error whenever you try to access it, due to extension security controls. It also can result in confusing, competing UIs for the user - the password manager intercepts the calls and presents its own UI, but if you cancel this you might get a browser dialog, which might itself hand over to a system dialog, all with their own style. Integration with the browser’s built-in webauthn support will provide a more seamless experience and allow the user to have passkeys from multiple password managers all available in one place.</p><p>Even the native browser APIs can have their quirks, though. For example, Safari requires the <a href="https://developer.mozilla.org/en-US/docs/Web/API/Web_Authentication_API" target="_blank" rel="noopener">webauthn API</a> call to be made in an event directly triggered by the user, like a click. However, the webauthn API requires a challenge, which requires an asynchronous HTTP call to the server to fetch it. So if you trigger this request on the user click, then call the webauthn API with the response, it will fail. Annoying.</p><p>The <a href="https://fidoalliance.org/specifications-credential-exchange-specifications/" target="_blank" rel="noopener">standard for people to export and import passkeys</a> is not yet complete, although it’s getting close. This will allow you to move between password managers without having to recreate your passkeys on every site. Once this is complete and widely adopted it will remove the risk of lock-in, which we believe is currently hampering adoption.</p><p>The good news is these problems are all solvable, and the ongoing work shows there is strong industry desire to do so. Despite the current minor issues, we believe the time is now right to start adopting passkeys, and we hope to see their continued success.</p></content> 935 </entry><entry> 936 <title>Get the best in email: how Domenic uses Fastmail</title> 937 <link rel='alternate' type='text/html' href='https://www.fastmail.com/blog/get-the-best-in-email-how-domenic-uses-fastmail/' /> 938 <id>https://www.fastmail.com/blog/get-the-best-in-email-how-domenic-uses-fastmail/</id> 939 <updated>2024-06-04T16:00:00Z</updated><author> 940 <name>The Fastmail Team</name> 941 </author><content xml:lang='en' type='html'><p>Fastmail makes it easy to organize and prioritize the messages that are most important to you. Domenic, a software developer, uses Fastmail personal domains, aliases, folders, labels, and more. Read on to learn how Domenic gets the most out of his Fastmail inbox and why he switched from Gmail.</p><h3 id="what-are-your-favorite-fastmail-features-and-what-purpose-do-you-use-them-for" tabindex="-1"><strong>What are your favorite Fastmail features, and what purpose do you use them for?</strong></h3><p>My favorite Fastmail feature is definitely personal domains. The ability to use my own domain names and create as many aliases for as many domains as I want makes my workflow so efficient.</p><p>I chose the Dark theme, and using the customizable colors and buttons is incredible. It’s great that Fastmail allows me to choose how I want my email client to look (and it does look beautiful 😊).</p><p>Email and calendaring are done well and feel very easy to use. It was a breeze to import all of my emails from my Gmail and set up a connection to send and receive emails via the Fastmail interface.</p><h3 id="how-do-you-maintain-your-workflow" tabindex="-1"><strong>How do you maintain your workflow?</strong></h3><p>I use Fastmail Rules to help categorize my incoming emails into a folder structure. The rules engine is powerful enough to perform all the options I need: move emails, delete emails, set follow-up flags, etc.</p><p>I pin all the important emails that I need to take further action on. Those emails appear at the top of my Inbox until I have dealt with them. This allows me to immediately see what items I need to follow up on, or messages that are awaiting a response.</p><hr><p>Fastmail is email for people created by people. With endless possibilities for personalizing your workflow and an expert support team available at any time, it’s easy to get better email today. <a href="/try-it/">Try Fastmail free for 30 days</a>!</p></content> 942 </entry> 943 </feed>