fal is a generative media cloud platform founded in 2021 by Burkay Gur and Gorkem Yurtseven. It provides access to more than 600 production-ready image, video, audio and 3D models through a single API, and supports building, deploying and training custom AI models at scale. The company operates globally and serves customers across generative AI infrastructure, software, creative tools and enterprise AI.
The platform rests on two main technical pillars: the fal Inference Engine, described as the fastest inference engine for generative media models and up to 10x faster than alternatives for models such as SDXL and Whisper; and a serverless GPU offering with on-demand compute priced from $1.2 per hour for premium GPUs. For larger workloads, fal runs dedicated compute clusters built on the latest NVIDIA hardware, including H100, H200 and B200 chips.
By scale, fal reports more than 1,000,000 developers on the platform, over 100m daily inference calls and 99.99% uptime, with end customers numbering in the hundreds of millions. Named customers include Canva, Perplexity, Quora's Poe and Adobe. Technical work at the company spans inference optimization, serverless GPU computing, machine learning infrastructure, and model training and deployment.
The company places emphasis on enterprise-grade security and maintains SOC 2 compliance.






