What you will learn
- Amazon FSx — four managed file systems and when each wins over EFS
- FSx for Windows File Server — Active Directory integration and SMB shares
- FSx for Lustre — high-performance computing and ML training storage
- FSx for NetApp ONTAP and OpenZFS — enterprise migrations
- AWS Storage Gateway — the bridge between on-premises and AWS cloud
- Three Storage Gateway types — File, Volume, and Tape
- AWS Snowball and Snowball Edge — physically moving large data to AWS
- AWS DataSync — automated online data transfer between storage systems
- AWS Transfer Family — SFTP, FTP, and FTPS into S3 and EFS
Why this matters
A Bangalore media company has a 500 TB video archive on-premises. Migrating it over the internet would take 185 days. Snowball devices ship to their office, they load the data locally at disk speed, ship the devices back, and the data appears in S3 in about a week. A financial services firm runs Windows servers and needs a shared drive that works with Active Directory — EFS is Linux-only and NFS. FSx for Windows File Server gives them a fully managed SMB share that integrates with their existing AD. Understanding which specialised storage service fits which situation separates engineers who know one tool from engineers who know the right tool.
Amazon FSx — Four Managed File Systems
EFS covers Linux NFS workloads. FSx covers everything else — Windows file shares, HPC workloads, and enterprise storage migrations. Four flavours:
FSx for Windows File Server:
A fully managed Windows file system. SMB protocol, Windows NTFS, Active Directory integration. Windows applications connect to it exactly like they connect to an on-premises Windows file server.
Use when: Servers are Windows-based Need to integrate with Active Directory for user permissions Migrating on-premises Windows file shares to AWS Need Windows features: shadow copies, DFS namespaces, Windows ACLs Supports: Multi-AZ for high availability Automated daily backups Encryption at rest with KMS Up to 2 GB/s throughput, millions of IOPSFSx for Lustre:
Lustre is a high-performance parallel file system used in supercomputing and large-scale ML. FSx for Lustre is managed Lustre.
Use when: Machine learning training reading large datasets (hundreds of GB per second needed) High-Performance Computing — simulations, video rendering, genomics Financial modelling with massive datasets Any workload needing sub-millisecond latency and extreme throughput Key capability — native S3 integration: Point FSx for Lustre at an S3 bucket Data loaded on demand as files are accessed Results written back to S3 automatically Your ML training job reads from Lustre at full speed, data comes from S3 Two deployment types: Scratch: temporary, no replication, cheaper — for short ML training jobs Persistent: replicated within one AZ, survives node failures — for long jobsFSx for NetApp ONTAP:
Fully managed ONTAP (NetApp's enterprise storage OS). For organisations already using NetApp on-premises who want to move to AWS without changing their storage layer.
Supports: NFS, SMB, iSCSI simultaneouslyCompatible with: Linux, Windows, macOS clients all at onceKey feature: SnapMirror — replicate ONTAP volumes from on-premises to FSxKey feature: Data deduplication and compression built inFSx for OpenZFS:
Managed ZFS file system. For Linux workloads that need ZFS-specific features — snapshots, cloning, compression, checksums.
Use when: migrating ZFS workloads from on-premises, need sub-millisecond latencyQuick selection:
| Need | Use |
|---|---|
| Linux shared storage, simple | EFS |
| Windows SMB, Active Directory | FSx for Windows |
| ML training, HPC, max performance | FSx for Lustre |
| Migrating NetApp on-premises | FSx for NetApp ONTAP |
| Migrating ZFS on-premises | FSx for OpenZFS |
AWS Storage Gateway — On-Premises Meets Cloud
Storage Gateway is a hybrid storage service. It runs as a virtual machine (or hardware appliance) on your on-premises servers. Applications talk to the Gateway using standard protocols. The Gateway stores data in AWS under the hood.
Three types — each for a different on-premises storage pattern:
File Gateway:
On-premises servers write files using NFS or SMBFile Gateway stores them as objects in S3Most recently accessed files cached locally for low latencyOlder files stored in S3 — retrieved on access Use case: On-premises application writes reports to a file share File Gateway transparently stores them in S3 Application sees a file share — AWS stores the filesVolume Gateway:
On-premises servers connect via iSCSI (appears as a block device)Two modes: Cached Volumes: primary data in S3, frequently accessed data cached locally Stored Volumes: all data stored on-premises, asynchronously backed up to S3 as EBS snapshots Use case: On-premises database needs block storage Automated daily backup to AWS without changing the application Restore entire volumes in AWS during a disaster recovery eventTape Gateway:
On-premises backup software writes to virtual tape drivesTape Gateway stores the tapes in S3 GlacierCompatible with leading backup software (Veeam, Backup Exec, NetBackup) Use case: Company already using tape-based backup Replace physical tape library with virtual tapes in Glacier Keep existing backup software — zero changesRememberStorage Gateway runs on-premises but stores data in AWS. File Gateway = NFS/SMB to S3. Volume Gateway = iSCSI block storage to S3/EBS snapshots. Tape Gateway = virtual tapes to Glacier.
AWS Snowball — Physical Data Transfer
When you have a lot of data and a slow internet connection, transferring it online takes too long.
200 TB at 100 Mbps internet connection:200,000 GB × 8 bits / 100 Mbps / 3600 / 24 = 185 days Same data via Snowball:AWS ships device → you load data at disk speed (1-10 GB/s) → ship back → ~1 week totalSnowball is a ruggedised physical device. AWS ships it to you. You connect it to your network, copy data onto it, ship it back. AWS loads the data into S3.
Snowball Edge — compute on the device:
Snowball Edge adds compute capability to the device. Run Lambda functions and EC2 instances directly on the Snowball Edge — useful when you need to process data at a remote site with no internet connectivity.
Use cases: Remote oil rig — collect IoT sensor data, process locally, sync when connected Military deployment — compute in the field with no internet Ship — process data while at sea, sync when dockingSnowball Edge types:
| Type | Storage | Compute | Use for |
|---|---|---|---|
| Storage Optimized | 80 TB | Basic | Pure data transfer |
| Compute Optimized | 28 TB | High (GPU optional) | Edge computing and ML |
OpsHub:
A GUI application to manage Snowball Edge devices. Drag and drop file transfers, manage EC2 instances running on the device, monitor device health.
RememberFor a one-time large data migration (TBs to PBs) → Snowball. For ongoing continuous data transfer → DataSync or Direct Connect. Snowball is for the initial bulk move, not for ongoing sync.
AWS DataSync — Automated Online Transfer
DataSync automates and accelerates moving data between:
On-premises storage ↔ S3, EFS, FSxS3 ↔ S3 (cross-region or cross-account)EFS ↔ EFSFSx ↔ FSxDataSync is not just a copy tool. It:
Verifies data integrity at source and destination — every file checksummedPreserves metadata — timestamps, permissions, ownershipSchedules transfers — run nightly, hourly, or continuouslyHandles network interruptions — resumes from where it stoppedUp to 10x faster than open-source tools like rsyncInstall a DataSync Agent on-premises:
Deploy DataSync Agent VM on your on-premises serverAgent connects to DataSync service in AWSConfigure: source (your NFS/SMB share) and destination (S3 bucket)Schedule: nightly at 1 AMDataSync transfers, verifies, and reports resultsDataSync vs Storage Gateway:
| DataSync | Storage Gateway | |
|---|---|---|
| Type | One-time or scheduled transfers | Continuous ongoing access |
| Pattern | Move data to AWS | Use AWS storage as an extension of on-premises |
| Use for | Migrations, scheduled sync | Hybrid cloud storage |
AWS Transfer Family — SFTP Into S3 and EFS
Transfer Family provides managed SFTP, FTPS, and FTP endpoints that store files directly in S3 or EFS.
Your trading partner sends files via SFTP to your Transfer Family endpointFiles land directly in your S3 bucketLambda triggered → processes the file immediatelyNo SFTP server to manage. No EC2 running an FTP daemon. Fully managed.
Supports: custom domains, Active Directory authentication, CloudWatch logging.
Use cases:
- Receiving files from partners who use SFTP
- Sharing files with customers via SFTP
- Regulatory file submissions via SFTP
Hands-on Lab — Storage Gateway File Gateway and DataSync
Step 1 — Explore FSx options in the console
AWS Console → FSx → Create file systemSee the four options: Windows, Lustre, NetApp ONTAP, OpenZFSClick Windows → review requirements: Active Directory domain, VPC, subnetsDo not create — just review the configuration optionsClick Lustre → review: S3 bucket integration, scratch vs persistentDo not create — review onlyStep 2 — Create a Storage Gateway (File Gateway)
Storage Gateway → Create gatewayGateway type: Amazon S3 File GatewayHost platform: Amazon EC2 (for testing — in production use on-premises VM)Launch EC2 instance using the provided AMI After instance launches:Storage Gateway → Select gateway → Service endpoint: Publicly accessibleActivate the gateway using the gateway IP Create S3 file share:File shares → Create file share → NFSS3 bucket: your existing bucketClient access: your VPC CIDRCreateStep 3 — Mount and test the File Gateway
On an EC2 instance in the same VPC:## Mount the File Gateway NFS sharesudo mount -t nfs \ -o nolock,hard \ GATEWAY-IP:/BUCKET-NAME \ /mnt/s3gateway ## Write a file through the gatewayecho "Stored via Storage Gateway" | sudo tee /mnt/s3gateway/test.txt ## Check S3 — the file appears as an objectaws s3 ls s3://your-bucket/ --region ap-south-1## test.txt should appearStep 4 — Set up DataSync (console walkthrough)
DataSync → Create agent (requires on-premises VM — for demo, review the setup) DataSync → Create task (review options):Source: your NFS share or S3 bucketDestination: S3, EFS, or FSxSchedule: daily at 2 AMOptions: verify data, preserve metadata, bandwidth limit In production this automates your nightly data sync to AWS.Step 5 — Cleanup
Storage Gateway → your gateway → Delete gatewayTerminate the EC2 instance used for the gatewayRemove the NFS mount: sudo umount /mnt/s3gatewayCommon Mistakes to Avoid
Common MistakeChoosing Snowball for ongoing data sync. Snowball is for bulk one-time migration. After the initial transfer, use DataSync for ongoing sync or Direct Connect for continuous connectivity. Snowball has a turnaround time of days — it cannot provide near-real-time sync.
Common MistakeUsing EFS when you need FSx for Windows. EFS uses NFS protocol and only works with Linux. Windows servers need SMB protocol. If you try to mount EFS on a Windows server it will not work. Windows shared storage needs FSx for Windows File Server.
TipFSx for Lustre integrated with S3 is one of the most powerful combinations for ML training. Your training data lives in S3 (cheap, durable). FSx for Lustre reads from S3 on demand at high speed. Your training job reads from Lustre at full throughput. When training finishes, results are written back to S3 automatically. You pay for Lustre only while the training job runs.