torch-mlir/include/npcomp/E2E/E2E.h

//===------------------------------------------------------------*- C++ -*-===//
//
// This file is licensed under the Apache License v2.0 with LLVM Exceptions.
// See https://llvm.org/LICENSE.txt for license information.
// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
//
//===----------------------------------------------------------------------===//

#ifndef NPCOMP_E2E_E2E_H
#define NPCOMP_E2E_E2E_H

#include "mlir/Pass/Pass.h"
#include "mlir/Pass/PassManager.h"

namespace mlir {
namespace NPCOMP {

/// Registers all E2E passes.
void registerE2EPasses();

// Look in createE2ELoweringPipeline for more information about how these
// passes fit together.
//
// Pass summaries are in Passes.td.

std::unique_ptr<OperationPass<FuncOp>> createLowerBroadcastToToLoopsPass();

std::unique_ptr<OperationPass<FuncOp>>
createLowerLinalgOnTensorToLinalgOnMemrefPass();

std::unique_ptr<OperationPass<ModuleOp>>
createLowerConstantTensorsToMemrefsPass();

std::unique_ptr<OperationPass<FuncOp>> createResolveShapeOfOpsPass();

std::unique_ptr<OperationPass<FuncOp>> createResolveTensorLoadStoreOpsPass();

std::unique_ptr<OperationPass<FuncOp>> createLowerLinalgLoopDimOpsPass();

std::unique_ptr<OperationPass<FuncOp>> createLowerRankedShapesPass();

std::unique_ptr<OperationPass<ModuleOp>> createLowerToNpcomprtABIPass();

std::unique_ptr<OperationPass<FuncOp>> createLowerAllocMemRefOpsPass();

std::unique_ptr<OperationPass<ModuleOp>> createLowerToLLVMPass();

void createLowerToHybridTensorMemRefPipeline(OpPassManager &pm);

struct E2ELoweringPipelineOptions
    : public PassPipelineOptions<E2ELoweringPipelineOptions> {
  // If this option is true, then perform optimizations.
  // If this option is false, only do the bare minimum for correctness.
  Option<bool> optimize{*this, "optimize", llvm::cl::desc("Do optimizations."),
                        llvm::cl::init(false)};
};

// The main pipeline that encapsulates the full E2E lowering.
void createE2ELoweringPipeline(OpPassManager &pm,
                               const E2ELoweringPipelineOptions &options);

} // namespace NPCOMP
} // namespace mlir

#endif // NPCOMP_E2E_E2E_H
Initial TCF/TCP E2E seed. Very much WIP. This is enough to get tcf.add down to approximately the "linalg.generic on buffers" level of abstraction. (but there are nuances) 2020-05-07 09:41:54 +08:00			`//===------------------------------------------------------------- C++ --===//`
			`//`
			`// This file is licensed under the Apache License v2.0 with LLVM Exceptions.`
			`// See https://llvm.org/LICENSE.txt for license information.`
			`// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception`
			`//`
			`//===----------------------------------------------------------------------===//`

			`#ifndef NPCOMP_E2E_E2E_H`
			`#define NPCOMP_E2E_E2E_H`

			`#include "mlir/Pass/Pass.h"`
			`#include "mlir/Pass/PassManager.h"`

			`namespace mlir {`
			`namespace NPCOMP {`

Bump submodule versions. * llvm-project: b5924a8e27536d19dd5c4d302db29fb6163d5faa * mhlo: 848ca244d20f045b7921da55a98a04d95ef94f0e * Multiple breakages that need to be fixed. Fixes: * Refactor dialect registration * Remove all kindof methods (Casting functionality has been added upstream and is implicitly available, see https://llvm.discourse.group/t/removing-kinds-from-attributes-and-types/1547.) * Update dialect registration to comply with https://reviews.llvm.org/D85495. * Remove type kinds and update some changed dialect signatures. * Upgrade ATen dialect to match upstream needs. * Move dialect registration to tablegen. * Register the ListType in tablegen. * Change dialect initialization signature. * Use TypeSwitch in MlirIr location printer. * Remove global registry depends from npcomp-opt. * Change LowerToLLVM to pass an MLIRContext vs an LLVMDialect for type creation. * Remove dep on MLIREDSCInterface that is removed upstream. * Thread through the DialectRegistry for opt and python-like tools. * Modernize pass registration (This was forced because the GEN_PASS_REGISTRATION code now generates inline functions vs literal pass registration statements) Co-authored-by: Marius Brehler <marius.brehler@iml.fraunhofer.de> 2020-08-28 06:09:10 +08:00			`/// Registers all E2E passes.`
			`void registerE2EPasses();`

Lower to the upstream memref ABI. Specifically, we use unranked memrefs which get passed as a fixed-size set of arguments/returns. One big caveat about this is that returning results isn't going to work. See TODO in LowerTensorLoadOp. This is far from enough runtime-wise, but it starts to demarcate a plausible layering. Notice for example how this removes the runtime-dependence from LowerRankedShapes. Eventually, we want to have an `npcomp_rt` or `npcomp_hal` dialect with its own set of runtime types that will supercede this. See comments in LowerTensorLoadOp for more direction about where this is going to evolve. 2020-05-16 07:33:01 +08:00			`// Look in createE2ELoweringPipeline for more information about how these`
			`// passes fit together.`
			`//`
			`// Pass summaries are in Passes.td.`

Initial TCF/TCP E2E seed. Very much WIP. This is enough to get tcf.add down to approximately the "linalg.generic on buffers" level of abstraction. (but there are nuances) 2020-05-07 09:41:54 +08:00			`std::unique_ptr<OperationPass<FuncOp>> createLowerBroadcastToToLoopsPass();`

			`std::unique_ptr<OperationPass<FuncOp>>`
			`createLowerLinalgOnTensorToLinalgOnMemrefPass();`

npcomprt: add support for constants - create tcp.global + tcp.get_global_memref - create npcomprt.global + npcomprt.get_global - LLVM lowering for new npcomprt ops - Runtime: - GlobalDescriptor struct emitted by LLVM lowering - implement __npcomp_compiler_rt_get_global Also, - cleanly isolate all runtime data structure definitions shared by the compiler and runtime into lib/runtime/CompilerDataStructures.h 2020-07-11 08:31:24 +08:00			`std::unique_ptr<OperationPass<ModuleOp>>`
			`createLowerConstantTensorsToMemrefsPass();`

Add script tools/format_source.sh and run it on all python and c++ sources. 2020-06-14 05:53:54 +08:00			`std::unique_ptr<OperationPass<FuncOp>> createResolveShapeOfOpsPass();`
Initial TCF/TCP E2E seed. Very much WIP. This is enough to get tcf.add down to approximately the "linalg.generic on buffers" level of abstraction. (but there are nuances) 2020-05-07 09:41:54 +08:00
"Finish" tensor -> memref conversion. There's a lot of details to flesh out here, but the basic approach seems promising (see comments in createE2ELoweringPipeline). This approach will be put to the test when we try to do our first fusions since that tickles some of the nasty phase ordering issues involved here. But we're not there yet. 2020-05-12 05:24:57 +08:00			`std::unique_ptr<OperationPass<FuncOp>> createResolveTensorLoadStoreOpsPass();`

Add lowering from linalg to loops. This also adds a small pass to clean up the `dim` ops that linalg introduces. For now, it only has a trivial pattern that looks for a `tcp.alloc_memref(%shape)` op to get the shape as we currently have an invariant that all memrefs are the result of such ops. But eventually this will need to look through view ops and any other shape-ish stuff that linalg introduces as it lowers to loops, along with any slicing ops introduced by buffer allocation. 2020-05-12 09:50:51 +08:00			`std::unique_ptr<OperationPass<FuncOp>> createLowerLinalgLoopDimOpsPass();`

Lower !shape.shape to SSA values. This uses an approach inspired by what is done in IREE. See comments on LowerRankedShapes.cpp for how it works. The basic gist is that we have an op that creates a !shape.shape from a set of SSA values representing the extents, and then iteratively replace any op producing a !shape.shape with instances of that op. 2020-05-12 12:22:40 +08:00			`std::unique_ptr<OperationPass<FuncOp>> createLowerRankedShapesPass();`

Rework e2e flow to use new "npcomprt" This ~totally reworks the existing "runtime" stuff to be more principled and usable, such as from Python. It's still not fully production-quality, mainly in the department of memory management (e.g. it currently leaks memory; we need to figure out "who frees memrefs" + the analysis and transformation needed to do that (maybe use upstream buffer allocation pass?)). The user API is in include/npcomp/runtime/UserAPI.h, though include/npcomp/JITRuntime/JITModule.h is a friendlier wrapper. The stuff under {include,lib}/runtime is totally firewalled from the compiler and tiny (<6kB, though no attention has gone into optimizing that size). For example, we don't link in libSupport into the runtime, instead having our own bare bones replacements for basics like ArrayRef (the JITRuntime helps with bridging that gap, since it can depend on all common LLVM utilities). The overall features of npcomprt is that it exposes a module that with multiple function entry points. Each function has arguments and results that are tensor-valued, and npcomprt::Tensor is the runtime type that is used to interact with that (and a npcomprt::Ref<T> reference-counting wrapper is provided to wrap npcomprt::Tensor in the common case). From an implementation perspective, an npcomprt module at the LLVM/object/binary level exposes a single module descriptor struct that has pointers to other metadata (currently just a list of function metadata descriptors). All interactions with the npcomp runtime are keyed off of that module descriptor, including function lookups and dispatching. This is done to dodge platform ABI issues and also allow enough reflection to e.g. verify provided arguments. Most of the compiler-side work here was in LowerToNpcomprtABI and LowerToLLVM. Also, - Rename npcomp_rt/NpcompRt to npcomprt/Npcomprt; it was getting annoying to type the underscores/caps. - misc improvements to bash_helpers.sh 2020-07-09 08:15:40 +08:00			`std::unique_ptr<OperationPass<ModuleOp>> createLowerToNpcomprtABIPass();`
Lower to the upstream memref ABI. Specifically, we use unranked memrefs which get passed as a fixed-size set of arguments/returns. One big caveat about this is that returning results isn't going to work. See TODO in LowerTensorLoadOp. This is far from enough runtime-wise, but it starts to demarcate a plausible layering. Notice for example how this removes the runtime-dependence from LowerRankedShapes. Eventually, we want to have an `npcomp_rt` or `npcomp_hal` dialect with its own set of runtime types that will supercede this. See comments in LowerTensorLoadOp for more direction about where this is going to evolve. 2020-05-16 07:33:01 +08:00
Lower tcp.alloc_memref ops to tcp.get_extent + std.alloc. - tcp.get_extent will be liminated while lowering shapes - std.alloc is supported by the upstream LLVM lowering. 2020-05-19 03:50:16 +08:00			`std::unique_ptr<OperationPass<FuncOp>> createLowerAllocMemRefOpsPass();`

Lower to LLVM dialect. With this commit, we finish conversion to LLVM dialect, and should be ready for subsequent commits to convert to an LLVM module and let LLVM codegen to native machine code. This required a custom "lower to LLVM" pass to support lowering tcp.abort_if to a runtime call. In the future, this pass will grow to do type conversions for our own runtime types as we add those. 2020-05-21 09:48:53 +08:00			`std::unique_ptr<OperationPass<ModuleOp>> createLowerToLLVMPass();`

Initial TCF/TCP E2E seed. Very much WIP. This is enough to get tcf.add down to approximately the "linalg.generic on buffers" level of abstraction. (but there are nuances) 2020-05-07 09:41:54 +08:00			`void createLowerToHybridTensorMemRefPipeline(OpPassManager &pm);`

Add ability to run without optimizations. The default is to only do the bare minimum needed for correctness, since that stresses the layering of the system maximally. 2020-06-02 10:30:13 +08:00			`struct E2ELoweringPipelineOptions`
			`: public PassPipelineOptions<E2ELoweringPipelineOptions> {`
			`// If this option is true, then perform optimizations.`
			`// If this option is false, only do the bare minimum for correctness.`
			`Option<bool> optimize{*this, "optimize", llvm::cl::desc("Do optimizations."),`
			`llvm::cl::init(false)};`
			`};`

Initial TCF/TCP E2E seed. Very much WIP. This is enough to get tcf.add down to approximately the "linalg.generic on buffers" level of abstraction. (but there are nuances) 2020-05-07 09:41:54 +08:00			`// The main pipeline that encapsulates the full E2E lowering.`
Add ability to run without optimizations. The default is to only do the bare minimum needed for correctness, since that stresses the layering of the system maximally. 2020-06-02 10:30:13 +08:00			`void createE2ELoweringPipeline(OpPassManager &pm,`
			`const E2ELoweringPipelineOptions &options);`
Initial TCF/TCP E2E seed. Very much WIP. This is enough to get tcf.add down to approximately the "linalg.generic on buffers" level of abstraction. (but there are nuances) 2020-05-07 09:41:54 +08:00
			`} // namespace NPCOMP`
			`} // namespace mlir`

			`#endif // NPCOMP_E2E_E2E_H`