FuryCore.ai

Proofs first, then silicon.

An AI computer on one chip, for training and inference. RISC-V cores run Linux and orchestrate FuryCore, an accelerator that puts the algorithms of modern AI models into silicon. Every block starts as a formal model; AI agents build against it, and proofs decide what ships.

  • Pre-seed
  • FPGA prototyping on AWS F2
  • Building the founding team

Built for Train and infer Language Vision Generation Speech Robotics

Yet another AI accelerator chip?

Every accelerator makes matrix multiplication fast. FuryCore does too.

But models spend their time elsewhere as well: recurrent state, expert routing, attention variants, gating, sampling loops and audio front-ends. On other chips that work runs on bolt-on units or on the host, with a round trip through memory between steps.

FuryCore gives each of those operators a verified tile of its own, and runs everything else on programmable compute.

Where the field sits

Each class below is built around a fast matrix-multiply primitive. FuryCore starts from the operators themselves.

  • GEMM engines and systolic arrays GPUs, TPUs
  • In-memory compute Axelera, d-Matrix
  • Dataflow SambaNova, Groq
  • RISC-V plus tensor units Tenstorrent, SemiDynamics
  • Wafer-scale Cerebras
  • FuryCore The operators themselves, as verified tiles, next to a matmul array and programmable compute.

Three ideas, one chip

Operators in silicon

Recurrence, expert routing, attention variants, sampling and audio front-ends get tiles of their own, next to a fast matmul and convolution array. Programmable compute runs the rest, so no model is locked out.

See the architecture

Formal first, AI-native

Each block starts as a formal model. AI agents write the hardware against it, and proofs gate every change before it merges.

See the method

An AI computer on one chip

RISC-V cores run Linux and real-time tasks and orchestrate the accelerator, for training and inference, from robots to on-prem racks.

See what it runs

An AI computer on one chip

One SoC. RISC-V harts run the operating system and the real-time loop, and FuryCore, the accelerator, sits at the centre of the die.

FuryCore SoC floorplan A die with HBM interfaces on both sides. Along the top: RISC-V Linux harts, RISC-V RTOS harts and programmable compute. In the centre: the FuryCore tile array. Along the bottom: on-chip SRAM, and PCIe and Ethernet. HBM HBM Linux harts RTOS harts Programmable compute FuryCore tile array SRAM PCIe Ethernet
Linux harts
The control plane: the Rust runtime, model loading, networking.
RTOS harts
The real-time plane. They drive FuryCore's command queues token by token, and in a robot they close the control loop.
FuryCore
The accelerator at the core of the die: operator tiles next to a fast matmul and convolution array.
Programmable compute
The generic path. Any operator without a tile runs here: slower, and still correct.
Memory and I/O
HBM streams the weights. On-chip SRAM holds activations and recurrent state. PCIe and Ethernet connect chips and hosts.

FuryCore tiles

A representative set. Each tile is specified against our software reference and checked bit for bit.

  • Gated DeltaNet recurrence, as a streaming state machine
  • Mixture-of-experts routing
  • Attention: GQA, sliding-window, sparse, bidirectional and shared-prefix, with online softmax
  • Vision patch embedding with 2D positions
  • Convolution, for CNNs, U-Nets and audio front-ends
  • FFT and mel-spectrogram front-end
  • Diffusion and flow sampler step, so denoising loops stay on chip
  • Token sampling, top-k and top-p
  • RMSNorm and LayerNorm
  • K-quant dequantization on the memory read path
  • Streaming GEMV and GEMM array

Train what you run

Backward passes and optimizer steps get tiles of their own, specified and gated like every other block. Gradients all-reduce across chips over the same relay that shards inference.

Design target Our software reference trains small transformers from scratch today.

  • Fuse on chip, stream only weights. Operators hand results to each other on chip instead of writing them back to memory between steps.
  • Promote what matters. An operator that new models lean on moves from programmable compute to a tile of its own in the next bitstream or tape-out.
  • One command interface. On AWS F2 the host drives FuryCore's command queue over PCIe today. On the SoC the RISC-V harts drive the same interface.

We have built a CPU this way before: riski5, a RISC-V core written in Clash, boots Linux on FPGA hardware. The SoC needs new, larger harts; riski5 shows the method works.

What it trains and runs

Aim Transformers and the simpler architectures before them: train and run them on one chip.

Hardware support for each family below is a design target from day one. Our software reference already runs much of it end to end, so every block has a bit-exact target to meet.

Reference runs today Reference in progress Next

FamilyTarget modelsServed bySoftware reference
Language Qwen3.8 (27B hybrid), Qwen3.8-Flash-Next (125B MoE), Gemma 4 (31B dense and 26B-A4B MoE), Qwen3.6-35B-A3B, diffusion LMs. Larger MoE models such as GLM-5.3 follow as memory allows. Recurrence and expert-routing tiles, sliding and sparse attention, dequantization on read. Reference runs today
Vision Recognition, OCR, and the vision towers in front of Gemma 4 and Qwen3.8. Patch embedding, bidirectional attention, convolution. Reference runs today
Image and video generation Diffusion and flow models: Stable Diffusion-class U-Nets, DiT-class image and video models. The sampler step on chip, convolution and attention tiles. Image: reference in progressVideo: next
Speech and audio Recognition, Whisper- and Parakeet-class. Generation, Piper-class text to speech. FFT and mel front-end, convolution, attention. Reference in progress
Decisions and encoders Jev-style typed decisions: Laya, a BERT-style ModernBERT cross-encoder, and Cloudflare's Clef on a Qwen backbone. BERT-style encoders and embeddings. One pass, no token loop: the state is computed once and every question branches off it. Shared-prefix attention, encoder attention, scoring heads. Embeddings: reference runs todayDecision models: next
Robotics Streaming 3D mapping, LingBot-Map-class: video in, camera poses and dense depth out. VLA policies: vision and language in, actions out. Real time on a power budget a robot can carry, with the RTOS harts closing the control loop. 3D mapping: reference runs todayVLA: next
Classic and small models CNNs, MLPs and RNNs. The matmul and convolution array, plus programmable compute. Next
Training Backward passes, optimizer steps, gradient all-reduce across chips. Training tiles, gated like every other block. Reference trains small transformersHardware: design target

On-prem racks

SoCs scaled out for serving and training under power, cost and sovereignty limits.

Robots and embedded systems

One SoC per robot or device: perception, mapping, speech, VLA policies and decisions on a power budget.

Bring-up starts small. A 0.6B model's weights already stream through the simulated core, byte for byte. Larger models follow, one memory step at a time.

Models load from FuryModel, our model file format: converted once from safetensors or GGUF, and laid out so weights stream straight into accelerator memory.

Formal first, AI-native

Formal models are our specification, not an afterthought.

Agents write the hardware. Proofs decide what ships.

  1. Specify Building

    Each block starts as an executable model in Haskell, with the properties it must hold. AI drafts the model; people review it.

  2. Generate In use

    AI agents write the Clash hardware and the Rust runtime against that model.

  3. Gate

    A change that fails a gate does not merge.

    • In use Property tests against the Haskell oracle, bit for bit
    • In use Co-simulation of the live gateware against our software reference, on real weights
    • Building Model checking and equivalence checking of the generated RTL
  4. Measure Building

    FPGA performance counters feed the next design loop.

  5. Carry Aim

    The same vendor-neutral cores move from FPGA to shuttle die to ASIC.

From Haskell to silicon

  1. Source
    • Haskell + Clash
    The only hardware source
  2. Generated
    • VHDL
    • Verilog
    • SystemVerilog
    With SVA and PSL properties alongside
  3. Consumed by
    • Vivado, for AWS F2
    • SymbiYosys and EQY, for proofs
    • An ASIC flow, for silicon
    Whichever HDL each tool takes

One typed source of truth: the proof, the FPGA image and the silicon all come from the same code, and nobody hand-translates RTL. Liquid Haskell checks properties at the source; equivalence checking closes the gap after code generation.

Aim Our aim: each new model architecture becomes checked hardware in weeks, not years.

Hardware in Haskell with Clash. Host and hart software in Rust. Every build reproducible with Nix.

Prototyping on AWS F2

We prototype on AWS EC2 F2 instances: AMD Virtex UltraScale+ VU47P FPGAs with high-bandwidth memory, rented by the hour.

Why F2

  • Decode is memory-bound. Each F2 FPGA brings 16 GiB of HBM at up to 460 GiB/s, plus 64 GiB of DDR4.
  • Vivado comes with the FPGA Developer AMI, so synthesis costs compute and nothing else.
  • One design scales from 1 to 2 to 8 FPGAs per instance.

What runs today

  • Our Nix-packaged F2 pipeline runs end to end up to the hardware run: Spot-priced synthesis, AFI creation, and a NixOS F2 runner image.
  • In simulation, the host software drives the live Clash gateware and streams real model weights through it.

Planned: model memory sets the instance

ModelsInstanceMemory
0.6B bring-up modelf2.6xlarge1 FPGA
Most vision, speech and image-generation models, under 16 GBf2.6xlarge1 FPGA
Gemma 4 26B-A4B, Qwen3.6-35B-A3B, quantizedf2.6xlarge to f2.12xlarge1 to 2 FPGAs
Qwen3.8-27B, about 62 GB working set, and Qwen3.8-Flash-Nextf2.48xlarge8 FPGAs, 128 GiB HBM
GLM-5.3 class, stretch goalf2.48xlarge512 GiB DDR4 as an expert tier

Every model will be checked against our software reference.

Growth path

  1. Short Spot runs on f2.6xlarge now.
  2. f2.12xlarge next, for FPGA-to-FPGA work.
  3. f2.48xlarge for the large models, sharded across 8 FPGAs.

Where we are

Technology readiness

TRL 2 to 3

Aiming at TRL 4 on F2 hardware.

Market readiness

Finding earlyvangelists

The technology exists; customer development has started.

Technology readiness, per subsystem

On the EU scale of technology readiness levels (TRL). A filled segment marks where a subsystem stands now; an outlined one marks its next target. The levels are our own assessment, tied to the evidence beside each.

Subsystem Now Scale 1 to 9 Evidence Next
Clash design flow to working FPGA hardware 4 riski5, our earlier RISC-V CPU written in Clash, boots Linux on FPGA hardware. Evidence for the method, not a FuryCore component. TRL 5 FuryCore designs on F2 hardware
Bit-exact testing against the software reference 3 A GEMV core matches its Haskell oracle under property tests and emits Verilog. TRL 4 Dequantization, GEMV and DeltaNet cores match on real weights
F2 platform 3 In simulation, the host software drives the live Clash gateware and streams real weights through it. Our AWS pipeline runs up to the hardware run. TRL 4 Hardware runs on F2
AI agents implementing hardware 3 Agents co-write the Clash and Rust today, behind the oracle tests. TRL 4 Agents gated by formal proofs
Operator cores 2 Specified against our software reference, which runs language, vision-tower and 3D-mapping models end to end. TRL 4 Cores validated on FPGA in the first funded phase
Formal specification and proof gate 2 Concept and tools chosen. Clash emits SVA and PSL properties next to the RTL. TRL 4 Operators and the queue protocol proven before merge
SoC: new RISC-V harts and integration 2 Concept. The harts are new development. TRL 4 Harts orchestrate FuryCore on one FPGA
Programmable compute and promotion loop 1 to 2 Architecture concept. TRL 3 First operators run on programmable compute
Training tiles 1 Concept. Our software reference trains small transformers from scratch, with hand-written backward passes. TRL 3 A backward tile checked against the reference
Silicon 1 to 2 Routes to multi-project wafers and partners identified. TRL 6 A shuttle die

TRL 2 to 3 today.

The EU TRL scale
  1. Basic principles observed
  2. Technology concept formulated
  3. Experimental proof of concept
  4. Technology validated in lab
  5. Technology validated in relevant environment (industrially relevant environment in the case of key enabling technologies)
  6. Technology demonstrated in relevant environment (industrially relevant environment in the case of key enabling technologies)
  7. System prototype demonstration in operational environment
  8. System complete and qualified
  9. Actual system proven in operational environment (competitive manufacturing in the case of key enabling technologies; or in space)

Roadmap

  1. F2 hardware TRL 4

    FuryCore designs run on AWS F2.

  2. SoC on an FPGA

    New RISC-V harts drive FuryCore on one FPGA.

  3. Partner workloads TRL 5

    Multi-FPGA systems run full models on design-partner workloads, and an embedded FPGA board runs perception and mapping in a partner's robot.

  4. Shuttle die TRL 6

    A slice of the SoC in silicon, on a multi-project wafer.

  5. Product ASIC TRL 7 and up

    The AI computer on one chip.

Aim Our aim: multi-thousand tokens per second on the ASIC, at a fraction of a GPU's power. A roofline target, not a measurement.

FPGA first and a shuttle die is how a small team crosses the TRL 5 to 7 valley of death without betting on a product tape-out.

Why Europe

Most AI in Europe runs on accelerators designed elsewhere. FuryCore builds accelerator IP here, for European racks and for Europe's robotics and industrial base.

Computing with digital sovereignty, without fearing that someone outside the EU pulls the plug.

The team

Mika Tammi

Founder, CTO and lead engineer

Systems engineer across the whole stack: SoC firmware and platform security, Linux and reproducible infrastructure, CPU design in Clash on FPGAs, and LLM inference. Builds infurer.ai, the Rust inference engine that serves as FuryCore's software reference.

Founding roles

Co-founder or founding engineer, by fit.

Commercial co-founder, CEO

Find the earlyvangelists, turn them into design partners and letters of intent, raise the rounds, and build the company.

Silicon and ASIC

Physical design, sign-off, DFT and tape-out. You take Clash-generated RTL through the ASIC flow, from the shuttle die to the product.

FPGA, RTL and SoC

The RISC-V harts and FuryCore on one fabric, timing closure on the VU47P, HBM, multi-FPGA bring-up. The RTL comes from Clash; you own what happens after it. Clash experience is not required, willingness to learn it is.

ML systems

Map new model architectures onto operator specs and the runtime, across language, vision, generation, speech, decision and robotics models, for training and inference. Python is not the requirement; turning Python reference code into production-grade Rust is.

We write hardware in Haskell with Clash, software in Rust, and build everything with Nix. If types, proofs and reproducible builds belong in hardware design to you, talk to us.

Company forming in Germany. Founder based in the German–Dutch border region; founding team remote-friendly across the EU, with regular in-person weeks.

We are forming the founding team now, ahead of a European deep-tech funding round this autumn.

Four ways in

Design partners

Running or training language, generation or speech models on your own hardware under power, cost or sovereignty limits?

Building robots or embedded systems that need perception, mapping or VLA policies on a power budget?

Tell us your model, latency, throughput and power envelope.

Email about your workload

Research partners

A formal methods, EDA, FPGA or ML systems group in Europe? Let's build a consortium.

Email about research

Angels and funders

Pre-seed. We talk with angels and deep-tech funds who back hardware early.

Email about investing

Or write to mikatammi@gmail.com.