{"id":915471,"date":"2023-01-29T22:29:33","date_gmt":"2023-01-30T06:29:33","guid":{"rendered":"https:\/\/www.microsoft.com\/en-us\/research\/"},"modified":"2023-04-25T16:54:45","modified_gmt":"2023-04-25T23:54:45","slug":"empowering-azure-storage-with-rdma","status":"publish","type":"msr-research-item","link":"https:\/\/www.microsoft.com\/en-us\/research\/publication\/empowering-azure-storage-with-rdma\/","title":{"rendered":"Empowering Azure Storage with RDMA"},"content":{"rendered":"
Given the wide adoption of disaggregated storage in public clouds, networking is the key to enabling high performance and high reliability in a cloud storage service. In Azure, we choose Remote Direct Memory Access (RDMA) as our transport and aim to enable it for both storage frontend traffic (between compute virtual machines and storage clusters) and backend traffic (within a storage cluster) to fully realize its benefits. As compute and storage clusters may be located in different datacenters within an Azure region, we need to support RDMA at regional scale.<\/p>\n
This work presents our experience in deploying intra-region RDMA to support storage workloads in Azure. The high complexity and heterogeneity of our infrastructure bring a series of new challenges, such as the problem of interoperability between different types of RDMA network interface cards. We have made several changes to our network infrastructure to address these challenges. Today, around 70% of traffic in Azure is RDMA and intra-region RDMA is supported in all Azure public regions. RDMA helps us achieve significant disk I\/O performance improvements and CPU core savings.<\/p>\n","protected":false},"excerpt":{"rendered":"
Given the wide adoption of disaggregated storage in public clouds, networking is the key to enabling high performance and high reliability in a cloud storage service. In Azure, we choose Remote Direct Memory Access (RDMA) as our transport and aim to enable it for both storage frontend traffic (between compute virtual machines and storage clusters) […]<\/p>\n","protected":false},"featured_media":0,"template":"","meta":{"msr-url-field":"","msr-podcast-episode":"","msrModifiedDate":"","msrModifiedDateEnabled":false,"ep_exclude_from_search":false,"_classifai_error":"","footnotes":""},"msr-content-type":[3],"msr-research-highlight":[],"research-area":[13547],"msr-publication-type":[193716],"msr-product-type":[],"msr-focus-area":[],"msr-platform":[],"msr-download-source":[],"msr-locale":[268875],"msr-post-option":[],"msr-field-of-study":[],"msr-conference":[],"msr-journal":[],"msr-impact-theme":[],"msr-pillar":[],"class_list":["post-915471","msr-research-item","type-msr-research-item","status-publish","hentry","msr-research-area-systems-and-networking","msr-locale-en_us"],"msr_publishername":"","msr_edition":"","msr_affiliation":"","msr_published_date":"2023-4-17","msr_host":"","msr_duration":"","msr_version":"","msr_speaker":"","msr_other_contributors":"","msr_booktitle":"","msr_pages_string":"","msr_chapter":"","msr_isbn":"","msr_journal":"","msr_volume":"","msr_number":"","msr_editors":"","msr_series":"","msr_issue":"","msr_organization":"USENIX","msr_how_published":"","msr_notes":"","msr_highlight_text":"","msr_release_tracker_id":"","msr_original_fields_of_study":"","msr_download_urls":"","msr_external_url":"","msr_secondary_video_url":"","msr_longbiography":"","msr_microsoftintellectualproperty":1,"msr_main_download":"","msr_publicationurl":"","msr_doi":"","msr_publication_uploader":[{"type":"url","viewUrl":"false","id":"false","title":"https:\/\/www.usenix.org\/system\/files\/nsdi23-bai.pdf","label_id":"243109","label":0}],"msr_related_uploader":[{"type":"url","viewUrl":"false","id":"false","title":"","label_id":"243112","label":0}],"msr_attachments":[],"msr-author-ordering":[{"type":"user_nicename","value":"Wei Bai","user_id":37035,"rest_url":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Wei Bai"},{"type":"text","value":"Shanim Sainul Abdeen","user_id":0,"rest_url":false},{"type":"text","value":"Ankit Agrawal","user_id":0,"rest_url":false},{"type":"text","value":"Krishan Kumar Attre","user_id":0,"rest_url":false},{"type":"user_nicename","value":"Victor Bahl","user_id":31167,"rest_url":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Victor Bahl"},{"type":"text","value":"Ameya Bhagat","user_id":0,"rest_url":false},{"type":"text","value":"Gowri Bhaskara","user_id":0,"rest_url":false},{"type":"text","value":"Tanya Brokhman","user_id":0,"rest_url":false},{"type":"text","value":"Lei Cao","user_id":0,"rest_url":false},{"type":"text","value":"Ahmad Cheema","user_id":0,"rest_url":false},{"type":"text","value":"Rebecca Chow","user_id":0,"rest_url":false},{"type":"text","value":"Jeff Cohen","user_id":0,"rest_url":false},{"type":"text","value":"Mahmoud Elhaddad","user_id":0,"rest_url":false},{"type":"text","value":"Vivek Ette","user_id":0,"rest_url":false},{"type":"text","value":"Igal Figlin","user_id":0,"rest_url":false},{"type":"user_nicename","value":"Daniel Firestone","user_id":35969,"rest_url":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Daniel Firestone"},{"type":"text","value":"Mathew George","user_id":0,"rest_url":false},{"type":"text","value":"Ilya German","user_id":0,"rest_url":false},{"type":"text","value":"Lakhmeet Ghai","user_id":0,"rest_url":false},{"type":"text","value":"Eric Green","user_id":0,"rest_url":false},{"type":"text","value":"Albert Greenberg","user_id":0,"rest_url":false},{"type":"text","value":"Manish Gupta","user_id":0,"rest_url":false},{"type":"text","value":"Randy Haagens","user_id":0,"rest_url":false},{"type":"text","value":"Matthew Hendel","user_id":0,"rest_url":false},{"type":"text","value":"Ridwan Howlader","user_id":0,"rest_url":false},{"type":"text","value":"Neetha John","user_id":0,"rest_url":false},{"type":"text","value":"Julia Johnstone","user_id":0,"rest_url":false},{"type":"text","value":"Tom Jolly","user_id":0,"rest_url":false},{"type":"text","value":"Greg Kramer","user_id":0,"rest_url":false},{"type":"text","value":"David Kruse","user_id":0,"rest_url":false},{"type":"text","value":"Ankit Kumar","user_id":0,"rest_url":false},{"type":"text","value":"Erica Lan","user_id":0,"rest_url":false},{"type":"text","value":"Ivan Lee","user_id":0,"rest_url":false},{"type":"text","value":"Avi Levy","user_id":0,"rest_url":false},{"type":"text","value":"Marina Lipshteyn","user_id":0,"rest_url":false},{"type":"text","value":"Xin Liu","user_id":0,"rest_url":false},{"type":"text","value":"Chen Liu","user_id":0,"rest_url":false},{"type":"text","value":"Guohan Lu","user_id":0,"rest_url":false},{"type":"text","value":"Yuemin Lu","user_id":0,"rest_url":false},{"type":"text","value":"Xiakun Lu","user_id":0,"rest_url":false},{"type":"text","value":"Vadim Makhervaks","user_id":0,"rest_url":false},{"type":"text","value":"Ulad Malashanka","user_id":0,"rest_url":false},{"type":"user_nicename","value":"Dave Maltz","user_id":31648,"rest_url":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Dave Maltz"},{"type":"user_nicename","value":"Ilias Marinos","user_id":39684,"rest_url":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Ilias Marinos"},{"type":"text","value":"Rohan Mehta","user_id":0,"rest_url":false},{"type":"text","value":"Sharda Murthi","user_id":0,"rest_url":false},{"type":"text","value":"Anup Namdhari","user_id":0,"rest_url":false},{"type":"text","value":"Aaron Ogus","user_id":0,"rest_url":false},{"type":"user_nicename","value":"Jitu Padhye","user_id":33179,"rest_url":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Jitu Padhye"},{"type":"text","value":"Madhav Pandya","user_id":0,"rest_url":false},{"type":"text","value":"Douglas Phillips","user_id":0,"rest_url":false},{"type":"text","value":"Adrian Power","user_id":0,"rest_url":false},{"type":"text","value":"Suraj Puri","user_id":0,"rest_url":false},{"type":"text","value":"Shachar Raindel","user_id":0,"rest_url":false},{"type":"text","value":"Jordan Rhee","user_id":0,"rest_url":false},{"type":"text","value":"Anthony Russo","user_id":0,"rest_url":false},{"type":"text","value":"Maneesh Sah","user_id":0,"rest_url":false},{"type":"text","value":"Ali Sheriff","user_id":0,"rest_url":false},{"type":"text","value":"Chris Sparacino","user_id":0,"rest_url":false},{"type":"text","value":"Ashutosh Srivastava","user_id":0,"rest_url":false},{"type":"text","value":"Weixiang Sun","user_id":0,"rest_url":false},{"type":"text","value":"Nick Swanson","user_id":0,"rest_url":false},{"type":"text","value":"Fuhou Tian","user_id":0,"rest_url":false},{"type":"text","value":"Lukasz Tomczyk","user_id":0,"rest_url":false},{"type":"text","value":"Vamsi Vadlamuri","user_id":0,"rest_url":false},{"type":"user_nicename","value":"Alec Wolman","user_id":30925,"rest_url":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Alec Wolman"},{"type":"text","value":"Ying Xie","user_id":0,"rest_url":false},{"type":"text","value":"Joyce Yom","user_id":0,"rest_url":false},{"type":"user_nicename","value":"Lihua Yuan","user_id":40955,"rest_url":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/microsoft-research\/v1\/researchers?person=Lihua Yuan"},{"type":"text","value":"Yanzhao Zhang","user_id":0,"rest_url":false},{"type":"text","value":"Brian Zill","user_id":0,"rest_url":false}],"msr_impact_theme":[],"msr_research_lab":[199565],"msr_event":[],"msr_group":[144899,715138],"msr_project":[860115,317750],"publication":[],"video":[],"download":[],"msr_publication_type":"inproceedings","related_content":{"projects":[{"ID":860115,"post_title":"Network Stack for Modern Cloud","post_name":"network-stack-for-modern-cloud","post_type":"msr-project","post_date":"2022-08-02 15:21:06","post_modified":"2023-03-05 01:11:10","post_status":"publish","permalink":"https:\/\/www.microsoft.com\/en-us\/research\/project\/network-stack-for-modern-cloud\/","post_excerpt":"As part of the Network Stack for 2030 initiative, we are rethinking the network stack, which was designed about 30 years ago, when the networks, and the applications they supported, looked very different.","_links":{"self":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-project\/860115"}]}},{"ID":317750,"post_title":"RDMA for Cloud Computing","post_name":"rdma-for-cloud-computing","post_type":"msr-project","post_date":"2016-11-07 15:54:05","post_modified":"2017-06-14 09:32:57","post_status":"publish","permalink":"https:\/\/www.microsoft.com\/en-us\/research\/project\/rdma-for-cloud-computing\/","post_excerpt":"In this project, we have introduced a series of technologies, including DCQCN congestion control and DSCP-based PFC, and addressed a set of challenges including PFC deadlock, RDMA transport livelock, PFC pause frame storm, slow-receiver symptom, to make RDMA scalable and safe, and to enable RDMA deployable in production at large scale. We currently are working on RDMA deadlock understanding and prevention, and RDMA support for future AI infrastructure. RDMA Congestion Control Modern datacenter applications demand…","_links":{"self":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-project\/317750"}]}}]},"_links":{"self":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-research-item\/915471"}],"collection":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-research-item"}],"about":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/types\/msr-research-item"}],"version-history":[{"count":4,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-research-item\/915471\/revisions"}],"predecessor-version":[{"id":929130,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-research-item\/915471\/revisions\/929130"}],"wp:attachment":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/media?parent=915471"}],"wp:term":[{"taxonomy":"msr-content-type","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-content-type?post=915471"},{"taxonomy":"msr-research-highlight","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-research-highlight?post=915471"},{"taxonomy":"msr-research-area","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/research-area?post=915471"},{"taxonomy":"msr-publication-type","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-publication-type?post=915471"},{"taxonomy":"msr-product-type","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-product-type?post=915471"},{"taxonomy":"msr-focus-area","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-focus-area?post=915471"},{"taxonomy":"msr-platform","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-platform?post=915471"},{"taxonomy":"msr-download-source","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-download-source?post=915471"},{"taxonomy":"msr-locale","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-locale?post=915471"},{"taxonomy":"msr-post-option","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-post-option?post=915471"},{"taxonomy":"msr-field-of-study","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-field-of-study?post=915471"},{"taxonomy":"msr-conference","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-conference?post=915471"},{"taxonomy":"msr-journal","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-journal?post=915471"},{"taxonomy":"msr-impact-theme","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-impact-theme?post=915471"},{"taxonomy":"msr-pillar","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-pillar?post=915471"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}