Personal project / Engine architecture

Tessera Engine

A game engine built around explicit responsibilities and measured decisions.

Tessera is the Windows game engine and editor I am developing in C++20, with Direct3D 12 and the Agility SDK. I write the ECS, job system, asset pipeline and renderer myself. This case study explains how those pieces fit together, why I chose their boundaries, and what I use to check that they work.

Decisions are measured rather than assumed. Subsystem tests, command-line probes, benchmark baselines and image comparisons accompany the implementation. The repository is private; the architecture, captures and measurements documented here provide a technical view of the work in progress.

  • C++20
  • Windows
  • Direct3D 12
  • Dear ImGui
  • Sharpmake
Editor overview with docking, viewport, hierarchy, inspector and Content Browser; this capture shows a scene with no cooked meshes loaded.
Editor overview with docking, viewport, hierarchy, inspector and Content Browser; this capture shows a scene with no cooked meshes loaded. Open original in a new tab

Overview

Small pieces, enforceable boundaries

Tessera takes its name from a piece of a mosaic. Each layer has one responsibility and a defined dependency direction. Simulation does not depend on rendering; the renderer builds a plan; the graphics backend translates that plan into GPU work. The editor consumes the same engine that it inspects.

My earlier Hibou Engine used OpenGL, EnTT and a WPF editor. Tessera carries those lessons into explicit GPU memory and synchronization, an ECS shaped around scheduling, and a C++ editor in the same process. The engine is a static library: the Agility SDK exports D3D12SDKVersion and D3D12SDKPath belong to the executable, so keeping that boundary explicit avoids a fragile DLL arrangement.

Dependencies

Layered architecture

The boundaries are enforced by folders as well as interfaces. Platform alone talks to SDL3; only RHI/D3D12/ includes Direct3D headers. Render decides what to draw without recording GPU commands. ECS, Jobs and Scene include neither Render/ nor RHI/; Assets and Project also remain independent of rendering.

Import libraries and compression belong exclusively to Tools/TesseraCooker. The runtime consumes cooked data without carrying authoring dependencies. Include scans check these rules, but compilation provides a stronger check: Tessera.NoD3D12.sln builds the engine and editor with a null backend and an SDL_Renderer interface. A leaked D3D12 type breaks that solution; its builds produce zero Dx*.obj objects.

Engine dependency layers and the isolated Direct3D 12 boundary.
Engine dependency layers and the isolated Direct3D 12 boundary. Open original in a new tab

Simulation

A custom ECS built for scheduling

I studied EnTT and Flecs, then built an ECS around three requirements: declared read/write access, deferred deterministic structural changes, and world snapshots. An Entity is an eight-byte index and generation; the generation detects stale references. Disk persistence uses a UUID, StableEntityId, rather than a runtime handle.

Entities with the same component set live in archetypes, stored as contiguous SoA columns in 64 KiB chunks. At two million entities, 64 KiB required 1,482 pages versus 5,955 at 16 KiB, with the same bandwidth. Sparse sets hold rare components without multiplying archetypes. Shared-component tables keep one value per distinct configuration, referenced by the chunk.

Each SystemDescriptor declares reads and writes. The scheduler derives a dependency DAG and runs non-conflicting systems together. With 32 systems, measured scheduling overhead was 0.68 µs per frame over 10,000 entities and 0.86 µs over 1,000,000: this cost follows systems, not entity count.

Fixed frame phases define when input, simulation, physics, events, structural playback, transforms and render extraction occur. Creation, destruction and component changes go through CommandBuffer, applied at one deterministic point. A development access auditor detects undeclared writes and is compiled out in Shipping. Typed events, observers and a dedicated hierarchy store complete the model; versioned transform propagation avoids recalculating stationary nodes.

Snapshot replay reproduced 120 identical audit records over 120 frames on the same machine and build. That is the verified scope of determinism, and the foundation for editor Play/Stop; cross-platform determinism is not promised.

struct Entity
{
    uint32_t Index = InvalidEntityIndex;
    uint32_t Generation = 0;
};
static_assert(sizeof(Entity) == 8,
    "Entity is a chunk column; its size is layout.");
Archetype storage split into chunks with contiguous component columns.
Archetype storage split into chunks with contiguous component columns. Open original in a new tab

Concurrency

Work stealing without blocking the frame

Each worker owns a Chase-Lev deque: it pushes and pops without a lock while idle workers steal from the other end. A worker waiting for a job executes other work on its current stack. The single-worker test checks that dependencies still progress. Threads outside the pool wait passively; allowing them to help would create two producers for a single-producer deque.

Jobs have AnyWorker, LongRunning or MainThread affinity. Separate counters track frame work and long-running work, so shader compilation cannot extend the frame wait. In the recorded comparison, the slowest frame during recompilation dropped from 803 ms to 18.75 ms. Per-worker FrameArena storage supports temporary allocation; ParallelFor over 1,048,576 elements produced the serial result with zero heap allocations.

Three concurrency defects surfaced during development, including one that only appeared with the editor window visible. A watchdog naming the stalled frame stage helped isolate it. Replacing batch stealing with element-by-element stealing moved that cost to the idle thread. A broad graph of real work measured 94.7% worker occupancy.

The cooker uses the same system: compressing one 2048² texture took 311.8 s with one thread and 9.7 s with 24. Tracy and an editor timeline expose CPU/GPU activity and work per worker. These are specific project measurements, not general throughput guarantees.

Profiling timeline showing jobs by worker.
Profiling timeline showing jobs by worker. Open original in a new tab

Persistence

One container, many chunks

Importing a source model means cooking it. A Tessera project stores engine formats and references the original through a named source root; it never copies the original into the project and rejects absolute source paths. This separates portable project data from the authoring files needed for a reimport.

The binary formats share a 64-byte header with magic, version, XXH3-128 content hash and chunk-table location. FourCC entries locate chunks aligned to 256 bytes. The runtime reads only the mip or LOD it needs. PROV records the source path and hash, cooker version and options; an unchanged-project scan took 102 ms versus 316 s for recooking. IDNT preserves the UUID across reimports.

A .tmesh contains an entire asset with internal submeshes: the buggy example turns 109 meshes into one file with 236 submeshes. SXFM stores per-piece transforms, allowing one wheel to be stored once and placed four times. In the measured collection, storage fell from 39.9 MB to 31.0 MB and draws from 342 to 229.

Optional chunks have defined absence semantics: missing SXFM means identity transforms. PROV was added without a version increase; SXFM raised the mesh version while older files still opened and rendered. This makes format evolution an explicit compatibility decision.

FormatResponsibility
.tmeshMesh asset, submeshes and LODs
.ttexBC-compressed textures and mips; BC6H for HDRI
.tmatOpenPBR material
.tscn / .tpfb / .tcellLevel, prefab and streaming cell
.tprojJSON project descriptor, outside the binary container
Cooked mesh layout and the planned meshlet extension.
Cooked mesh layout and the planned meshlet extension. Open original in a new tab

Planned extension

Preparing the path to mesh shaders

Mesh shaders are not implemented in the renderer. The existing pipeline deliberately leaves the path open: canonical root-signature slots are visible to all stages, pipeline creation uses the pipeline state stream, DXC recognizes ms and as at Shader Model 6.5 or later, and capability detection queries GPU support.

A test successfully created a mesh-shader pipeline against the canonical root signature. That checks compatibility, not a shipping rendering path. meshoptimizer is selected but not integrated. The plan is for the cooker to generate meshlets into a new .tmesh chunk; older files will retain the traditional vertex path when that chunk is absent. The container and pipeline contracts are what make that extension possible.

GPU boundary

The renderer plans; the backend executes

IRenderBackend exposes no Direct3D types. Resources use generation-checked handles such as Handle<TextureTag> and Handle<MeshTag> to detect references to released objects. Render/ builds a GPU-independent plan; RHI/D3D12/ executes it. This makes planning testable without a device.

The backend defaults to Enhanced Barriers with a legacy fallback, and uses bindless Shader Model 6.6 with descriptor-table fallback. Its root signature occupies 13 of 64 DWORDs. A separate copy queue handles uploads; VRAM-budgeted streaming falls back to smaller mips under pressure, instead of dropping the scene. Shader hot reload rebuilds affected pipelines.

Multithreaded recording is selected where the workload justifies it. At 23,099 draws, measured CPU frame time dropped from 207–229 ms to 82–91 ms, a 2.3–2.8× improvement. The ordinary small frame deliberately keeps a single command list. Extra scheduling and recording machinery should earn its cost.

The render plan crosses an abstract interface before the D3D12 backend executes it.
The render plan crosses an abstract interface before the D3D12 backend executes it. Open original in a new tab

Frame compilation

A render graph that produces a plan

A pass declares resource reads and writes, access types and GPU stages. The graph sees the whole frame and derives ordering and barriers from those declarations. It never records commands. The D3D12 executor reads the execution order and per-pass barriers, then records the corresponding work.

Compilation builds dependencies against the preceding last writer, rejects cycles with an empty plan, removes unused producers, selects queues, computes transient lifetimes and generates barriers where resource roles change. Exported resources and declared side effects preserve required work; simply importing a resource does not protect its producer from removal.

Use order and lifetime are distinct with optional compute: positions in the plan do not describe overlapping execution. The graph provides lifetime intervals; the heap owner chooses aliasing and requests aliasing barriers. My allocator packs by lifetime, the requirement that led me to reject the evaluated size-based D3D12 Memory Allocator approach. Eight deliberately staggered resources used 2,688 KB instead of 10,752 KB, a 4.00× reduction and the minimum for that test.

The supplied forward graph was exported from Sample.tscn on October 6, 2026. In that recorded configuration, 20 passes produce 15 barriers. Six disabled occlusion/ray-tracing passes only open empty scopes: a pass with no writes is not an unused producer and is not automatically removed. The highlighted SceneColor transition before ToneMapPass is inferred from its change from render target to shader read. Counts vary with the shading path and enabled effects.

Reset() clears frame data while retaining vector capacity; graph resource handles expire with it. Graphviz export and the editor panel expose the plan, while --rendergraph-self-test runs 242 checks without a GPU.

  1. Declare

    Import existing resources with their incoming state; create undefined transients; use AddPass to declare Read, Write or ReadWrite.

  2. Compile

    Export required end states, including present for the backbuffer. Compile() derives a dependency order, resource lifetimes and synchronization.

  3. Execute

    The backend consumes GetExecutionOrder() and GetBarriers() and emits GPU work. Planning remains independent of Direct3D 12.

Forward frame graph exported from Sample.tscn on October 6, 2026; the highlighted SceneColor transition is inferred by the graph.
Forward frame graph exported from Sample.tscn on October 6, 2026; the highlighted SceneColor transition is inferred by the graph. Open original in a new tab

Rendering

Lighting with a reference for comparison

The renderer implements OpenPBR coat, anisotropy and fuzz, checked with a furnace test; image-based lighting, procedural sky and HDRI; cascaded shadows and clustered light culling. In the recorded culling test, 125× more lights cost 1.39×. Forward and deferred paths are interchangeable, with SSAO, bloom, histogram auto-exposure and transparency.

Inline DXR RayQuery provides soft shadows, ambient occlusion and traced reflections, with reflections limited to deferred rendering. A spatial-temporal denoiser and adaptive Low/High/Ultra budgets control the real-time effects. A reference path tracer measures their error. The captures show the sample scene and diagnostics; their visible counters describe those captures rather than a benchmark for every scene.

Sponza sample scene rendered inside the Tessera editor.
Sponza sample scene rendered inside the Tessera editor. Open original in a new tab

Authoring

The editor uses the systems it exposes

Dear ImGui docking keeps the editor in C++, in the engine process. The Lobby opens a project without launching another application. The Content Browser has rendered thumbnails, a self-invalidating disk cache and background imports. Mesh/prefab windows provide GPU previews; the texture editor isolates channels and packs ORM. Asset properties expose the binary chunks directly.

Hierarchy, Inspector, ImGuizmo transforms, undo/redo, prefabs, tags and asset colours support scene work. The Material Graph is a read-only view derived from .tmat; its wiring is not editable. The graph widgets come from ImGuizmo’s GraphEditor.

Play, Pause, Stop and Simulate share the snapshot mechanism: Stop restores the captured world and verifies that its hash is identical. Authoring scenes use JSON for readable Git diffs and are cooked to binary for runtime. Profiling, Render Graph, ECS Diagnostics and Console make the running systems inspectable from the editor.

Asset editor with a GPU-rendered 3D preview.
Asset editor with a GPU-rendered 3D preview. Open original in a new tab

Tradeoffs

Dependencies with a defined role

I write the engine-specific policies and use libraries for bounded problems. SDL3 sits behind IPlatformBackend; DirectXMath supplies SIMD mathematics across the engine; DXC compiles Shader Model 6.x HLSL. Dear ImGui and ImGuizmo supply interface primitives rather than a second managed editor stack.

cgltf, ufbx and my OBJ importer remain inside the cooker. I chose open-source ufbx over the FBX SDK for its triangulation; mikktspace matches tangent conventions used by art tools. Among the evaluated compressors, AMD Compressonator gave the preferred BC7 quality. stb_image and tinyexr decode images/HDRI; xxHash identifies cooked content; nlohmann/json preserves key order for stable text diffs.

Tracy supplies CPU/GPU timelines, and WinPixEventRuntime names GPU passes in PIX. vcpkg pins dependency versions by manifest baseline. Before adding a dependency I review its component licence at the official source, with a preference for open source.

Validation

Builds and comparisons as design constraints

I returned to Sharpmake after CMake so one generator defines both IDE and command-line builds. Sharpmake and vcpkg produce two solutions in Debug, Development, Profile and Shipping: eight local builds, all treating warnings as errors. Post-build checks verify runtime DLLs and mirrored shader trees.

Headless subsystem tests include 4,883 ECS checks, 242 render-graph checks and 837 cooker checks. Command-line probes exercise individual features; a benchmark harness compares timings and counters against a baseline. Image diffs against the preceding version require zero different scene pixels for the refactor acceptance checks described here.

Those numbers come from the project records used for this case study. They document the validation method and measured scenarios, not a fresh execution of the private repository by this website. Together, the alternate build, replay audits and image comparisons check different failure modes that a successful normal build alone would miss.

In progress

What remains to build

Meshlets and mesh shaders are the next rendering steps. Jolt Physics, miniaudio, GameNetworkingSockets and Recast & Detour are selected dependencies, not integrated systems. An authoritative server for 16 players is a goal; replication-scoped snapshots and a precision measurement for worlds up to 200 km² are groundwork, but networking has not started.

Temporal upscaling, TAA and motion blur remain future work. Async compute is implemented, optional and disabled by default; its measured gain was about 2%, below the 5% threshold used to justify enabling it. Global illumination is not implemented. Windows is the only supported platform.

There are also unresolved lifecycle details: switching projects in the same process still accumulates meshes in GPU memory. I treat that as unfinished work. The point of these notes is to make both the implemented architecture and its present limits visible.

Contact

Québec City, QC, Canada / Available for opportunities

I'm open to opportunities in software development, tools programming and gameplay programming, as well as junior 3D artist roles. If my experience could be a good fit for your team, I'd be happy to connect through my social profiles.