StepFun
Chinese multimodal foundation model lab founded by a former Microsoft researcher
StepFun is a Chinese artificial intelligence company that develops multimodal foundation models, known as the Step series, covering text, vision, and video understanding and generation tasks. Founded by a researcher with a background at…
Definition
StepFun is a Chinese artificial intelligence company that develops multimodal foundation models, known as the Step series, covering text, vision, and video understanding and generation tasks. Founded by a researcher with a background at Microsoft, StepFun pursues large-scale multimodal model training as its primary technical bet, offering both API access and consumer-facing applications built on its models. It is generally grouped with the newer wave of well-funded Chinese AI labs formed specifically to compete on foundation-model research.
Overview
StepFun was founded by Jiang Daxin, a former Microsoft executive with a research background in natural language processing and search, who assembled a team to pursue large-scale multimodal foundation models as the core of the company's technical strategy. It emerged during the period when numerous well-funded Chinese AI labs were formed in quick succession, each betting on a particular technical angle, and StepFun's chosen angle has consistently been multimodality: building models that jointly reason over text, images, and video rather than treating each modality as a bolt-on feature. StepFun's Step-series models are trained using large-scale multimodal datasets that pair text with visual and, in some cases, video content, allowing the resulting models to perform tasks like image understanding, visual question answering, and video comprehension in addition to standard text generation. The company has published research describing scaling its multimodal architectures to very large parameter counts, following the general industry pattern of scaling both model size and training data to improve capability, though as with other labs its precise training details are only partially disclosed. StepFun has also released some smaller model checkpoints openly, giving outside researchers a way to inspect and build on its multimodal architecture choices without relying solely on its hosted API. Within the Chinese AI landscape, StepFun is often grouped with MiniMax as a multimodal-first company, in contrast to text-first labs like Zhipu AI, Baichuan AI, and Alibaba's Qwen team. StepFun differentiates itself by emphasizing large-scale multimodal pretraining as a research bet from the outset rather than adding multimodal capability to an already-established text model, and by its founder's specific research pedigree in search and language understanding from his Microsoft years. It is generally viewed as one of the more research-driven of the newer Chinese AI labs. In practice, StepFun's models are used by developers building applications that need to understand or generate content combining text and visual information, such as visual assistants, image-captioning tools, and multimodal chatbots. The company also offers consumer-facing products that let end users interact directly with its multimodal capabilities, and its API allows third-party developers to integrate Step-series models into their own applications. As with other emerging Chinese AI labs, the practical limitations for international adopters include sparse English-language documentation, a smaller developer ecosystem compared to established multimodal offerings from OpenAI or Google, and less mature third-party tooling. Organizations evaluating StepFun should also weigh the relative immaturity of the company compared to longer-established multimodal providers, since track record and long-term model support are harder to assess for a newer entrant still establishing its market position. Teams that need guaranteed long-term API stability may prefer a more established multimodal provider until StepFun's product and support commitments mature further.
Key Features
- Develops the Step series of multimodal foundation models
- Founded by Jiang Daxin, a former Microsoft research executive
- Focuses on joint text, image, and video understanding from the outset
- Trains large-scale multimodal architectures on paired text-visual data
- Offers both API access and consumer-facing multimodal applications
- Grouped among multimodal-first Chinese AI labs alongside MiniMax
- Publishes research on scaling multimodal model architectures
- Positions itself as a research-driven entrant in the Chinese AI market