torch-mlir

Commit Graph

Author	SHA1	Message	Date
Stella Laurenzo	3d74337be0	Add a torch.kernel_call op and associated predicates.	2020-09-29 15:10:38 -07:00
Stella Laurenzo	2c9ca79c89	Add boilerplate for Torch dialect.	2020-09-28 15:26:17 -07:00
Sean Silva	16c26ef57e	[RefE2E] Use upstream shape constraint conversion pass. Now that we upstreamed our pass, we can remove it. The final pass that landed upstream doesn't do the shape.assuming canonicalization to legalize that op away, so added a restricted-canonicalizer pass that allowed to run just shape dialect canonicalizations, which deletes the shape.assuming. The pass ended up kind of ugly. See the TODO's on it for some potential cleaner directions.	2020-09-28 09:34:44 -07:00
Sean Silva	6ea37cfed6	Bump llvm-project to 9ed1e5873c19eb817fb9e36d0262c7effee5d35e Date: Fri Sep 18 13:55:52 2020 -0700 - Update to linalg syntax - New generated builders are better. Custom builder for tcp.shaped_results is now redundant.	2020-09-28 09:34:44 -07:00
Sean Silva	f9b37c55b7	[RefE2E] Add support for unary ops exp and tanh This is fairly mechanical.	2020-09-24 18:41:30 -07:00
Sean Silva	c69e9fabc5	[RefE2E] Add support for "max". This cleans up the lowering pipeline to easily allow extending to multiple binary ops. It looks fairly repetitive at multiple levels, but I don't want to prematurely generalize. I think that in principle we could derive a large swatch of TCF + TCP from a single linalg-style specification. Another direction is to use an OpInterface (something like "buildLinalgGenericBody"). I'm keeping my eye on it. In a subsequent commit, I'll mechanically add a set of binary ops modeled off of the std arithmetic ops.	2020-09-22 18:38:32 -07:00
Sean Silva	7b7f35744b	[RefE2E] Add interesting control flow example. This also required adding a lowering for ForOp in our tensor->memref conversion.	2020-09-21 12:25:24 -07:00
Sean Silva	276f5b80ea	[RefE2E] Add assemblyFormat for TCF and TCP ops and tidy up.	2020-09-18 15:03:53 -07:00
Sean Silva	dc8afc9271	[RefE2E] Refactor how tcf.add is lowered. It was previously going through this awkward route that prematurely created linalg.generic ops, which was an annoying layering problem since we can't compute a shape transfer function for linalg.generic in the general case. Now we pass it through the same path as tcp.matmul, with the shape transfer function being defined for tcp.add. This also removed the need for TCPToLinalg (now deleted). The equivalent of that is happening in lower-shaped-results-to-memref. One interesting outcome of this: we're basically using linalg as a "Buffer TCP". We might want to look into using named structured ops for more of TCP, but that would be a big velocity hit since then any change to the ODS / verification for those ops would be a change to the upstream structured op ODS generator. After we have more experience defining this manually, we should re-evaluate rebasing TCP on generated named linalg ops.	2020-09-18 15:03:53 -07:00
Sean Silva	d8675f8ad2	[RefE2E] Add support for matmul. I'm pretty happy with how this turned out. It looks pretty much like it should -- one change at each layer. This particular op bottoms out on linalg which takes care of the rest. - Add tcf.matmul - Add tcp.matmul - Add TCF->TCP lowering - Add tcp.matmul shape transfer function (BypassShapes.cpp) - Add tcp.matmul -> linalg.matmul lowering (LowerShapedResultsToMemref.cpp) - Add support to LowerShapeConstraints for lowering the new shape.cstr_require This matmul op is pretty limited in its capabilities. There is no batching and no multidimensional contraction. Certainly more design work will be needed to find the right abstractions that aren't too general but also help to canonicalize many cases from frontends. This is mainly to show that adding a new op needn't be very "scary" once we have the e2e infra in place. Also, - this clears out some exploratory cruft from the TCF dialect now that this is starting to become real.	2020-09-18 11:31:01 -07:00
Sean Silva	62738d3641	[RefE2E] Fix nul-termination bug. I was seeing some of the error messages come out with some garbage at the end. This fixes it.	2020-09-18 11:31:01 -07:00
Sean Silva	75f57b461e	Totally rework RefE2E tensor to memref flow. (#42 ) This now gets the overall "RefE2E" compilation stack to a point that I'm fairly happy with. We simplify it by mostly embracing the "descriptor" view of the world. The overall flow is best understood by reading through the createE2ELoweringPipeline function in lib/E2E/E2E.cpp That function creates a pass pipeline that lowers from "TCF" (which is ~numpy level of abstraction) down to LLVM IR. A brief high-level summary of what happens there: 1. TCF to TCP conversion. This involves reifying error handling in the form of shape constraints. See test/Conversion/TCFToTCP/basic.mlir 2. Lowering shape constraints. This converts shape constraints into eager error-handling code. See test/E2E/lower-shape-constraints.mlir This pass will soon go upstream. Because this lowers to std.assert, some later passes like LowerToNpcomprtABI and LowerToLLVM are updated to properly plumb this through e2e. See test/npcomp-run-mlir/invalid-broadcast.mlir for an execution test that properly aborts in case of an error. 3. Lowering tensors to memrefs. This is done via a series of passes rather than an single mega conversion. Unlike the previous code that mixed in the npcomprt ABI stuff here, it's now a very clean "pure memref" conversion. See test/E2E/lower-*-to-memref.mlir and lib/E2E/TensorToMemref/ Most of the changes are concentrated here. 4. As part of the above, we use the upstream ConvertShapeToStandard for lowering shapes. 5. We lower linalg to loops and lower loops to CFG using upstream passes. 6. Rewrite the "ABI" boundaries of the program to npcomprt data structures (LowerToNpcomprtABI). This mainly affects ABI boundaries and how global tensor constants are represented. One of the major improvements in this commit is that now it's a very clean rewrite that just replaces memrefs on ABI boundaries with !npcomprt.tensor (before there was a get_extent function that is not needed). See test/E2E/lower-to-npcomprt-abi.mlir 7. Lower to LLVM with upstream mlir patterns + some patterns for the npcomprt lowerings. One aspect here that is still a remnant of a non-descriptor-based tensor to memref flow is the BypassShapes + LowerShapedResultsToMemref. BypassShapes wraps the "tensor compute" ops in a tcp.shaped_results (basically a "tie_shape" kind of op), and then LowerShapedResultsToMemref uses those annotations to allocate output buffers while lowering the "tensor compute ops". Note that there are very few "tensor compute" ops currently supported (tcp.add + tcp.broadcast_to), so we just hardcode them in both passes. Realistically, I expect this to go away as we fully embrace the descriptor-based approach for simplicity, so don't look too deep into it.	2020-09-16 17:31:40 -07:00
Stella Laurenzo	97d83f786a	Bump submodule versions. * llvm-project: b5924a8e27536d19dd5c4d302db29fb6163d5faa * mhlo: 848ca244d20f045b7921da55a98a04d95ef94f0e * Multiple breakages that need to be fixed. Fixes: * Refactor dialect registration * Remove all kindof methods (Casting functionality has been added upstream and is implicitly available, see https://llvm.discourse.group/t/removing-kinds-from-attributes-and-types/1547.) * Update dialect registration to comply with https://reviews.llvm.org/D85495. * Remove type kinds and update some changed dialect signatures. * Upgrade ATen dialect to match upstream needs. * Move dialect registration to tablegen. * Register the ListType in tablegen. * Change dialect initialization signature. * Use TypeSwitch in MlirIr location printer. * Remove global registry depends from npcomp-opt. * Change LowerToLLVM to pass an MLIRContext vs an LLVMDialect for type creation. * Remove dep on MLIREDSCInterface that is removed upstream. * Thread through the DialectRegistry for opt and python-like tools. * Modernize pass registration (This was forced because the GEN_PASS_REGISTRATION code now generates inline functions vs literal pass registration statements) Co-authored-by: Marius Brehler <marius.brehler@iml.fraunhofer.de>	2020-09-08 13:26:42 -07:00
stephenneuendorffer	bb668e6e26	Add ATen Dialect (#16 ) This patch adds a dialect intended to be used as a frontend dialect to facilitate lowering from "A Tensor Library" in torch/pytorch. This patch includes several passes that are useful in conjuction with the dialect: --aten-layer-name: Generates layer names for each operation, which are not present in the original pytorch. --aten-to-std: Lower the ATen dialect into standard dialect function calls. --return-elimination-pass: convert functions (primarily the toplevel function) to pass return values by reference. This simplifies pytorch integration. --aten-op-report: generate a textual report about the model --liveness-report Future patches will implement actual integration with the pytorch jit to intercept and generates MLIR in this dialect, then lower the resulting MLIR into function calls through aten-layer-name -> aten-to-std -> return-elimination -> std-to-llvm. The result would then jitted using the LLVM jit, linked against a runtime library which makes calls back into pytorch to implement all the layers. Co-authored-by: Jeff Fifield <jeff.fifield@xilinx.com> Co-authored-by: Jeff Fifield <jeff.fifield@xilinx.com>	2020-08-12 19:28:04 -07:00
stephenneuendorffer	5beaf4cc01	Fix build again (#14 ) The RuntimeShlib.so now lives in /lib.	2020-08-07 08:36:03 -07:00
Stella Laurenzo	38abe99805	Collapse python_native/ into python/. * These were separated originally for layering reasons that no longer apply. * Most of the python extension code is under lib/ with just the module setup in python/.	2020-08-03 17:46:34 -07:00
Stella Laurenzo	571c8b448a	Collapse different top level test directories into test/. * Uses local configs and unsupported annotation to disable optional tests. * This separation was just an artifact of having initial trouble getting lit setup.	2020-08-03 17:41:16 -07:00
Stella Laurenzo	fc484d1bd8	Rework reference shape lowering based on upstream shape dialect changes. * Primarily, the upstream shape dialect now uses tensor<?xindex> for non-erroring, immediate shape calculations (and will return this for shape_of of a tensor or memref). * In addition, upstream passes do not yet exist for fully lowering to standard ops, so the passes here need to be extended to handle this new convention. * This should be seen as an intermediate state, necessary to integrate a new LLVM version and needs more work and cleanup for generality. * There is a good deal of awkwardness in these conversions. The hope is that additional upstream work will yield better defined conversion paths once out of this intermediate state.	2020-08-03 13:43:49 -07:00
Sean Silva	3f3dcad871	Make input file to npcomp-run-mlir be positional. This makes command lines more succinct.	2020-07-13 16:02:19 -07:00
Sean Silva	df0d3fcaff	Consolidate LLVM definitions of runtime data structures. This required making module descriptors hold a FuncDescriptor* instead of a pointer to array of FuncDescriptors as it previously did, which is innocuous (just requires an llvm.bitcast after the llvm.mlir.addressof).	2020-07-10 17:50:55 -07:00
Sean Silva	e228aa4b11	npcomprt: add support for constants - create tcp.global + tcp.get_global_memref - create npcomprt.global + npcomprt.get_global - LLVM lowering for new npcomprt ops - Runtime: - GlobalDescriptor struct emitted by LLVM lowering - implement __npcomp_compiler_rt_get_global Also, - cleanly isolate all runtime data structure definitions shared by the compiler and runtime into lib/runtime/CompilerDataStructures.h	2020-07-10 17:31:24 -07:00
Stella Laurenzo	efbcf0aa44	Add NumpyPublicFunctionsToTensor pass. * Rewrites public function signatures to operate on tensors (vs ndarray). * Most of our backends presume immutable tensors at public function boundaries.	2020-07-08 22:51:54 -07:00
Stella Laurenzo	5ceb37c19b	Add NumpyToTCF conversion. * Just for numpy.add right now.	2020-07-08 21:03:57 -07:00
Sean Silva	f18014f60c	LowerRankedShapes: support shape.const_shape op. Also, the previous code had a special case for deleting this op when it had no uses. This is subsumed by the change in this commit since now shape.const_shape is properly lowered. With this change, the included test case with multiple serially dependent ops works! This specific issue was related to the scalar argument to that function. We needed to compute a broadcast of a scalar shape (which is a shape.const_shape) with another shape.	2020-07-08 20:12:40 -07:00
Sean Silva	b4f0cea8fa	Rework e2e flow to use new "npcomprt" This ~totally reworks the existing "runtime" stuff to be more principled and usable, such as from Python. It's still not fully production-quality, mainly in the department of memory management (e.g. it currently leaks memory; we need to figure out "who frees memrefs" + the analysis and transformation needed to do that (maybe use upstream buffer allocation pass?)). The user API is in include/npcomp/runtime/UserAPI.h, though include/npcomp/JITRuntime/JITModule.h is a friendlier wrapper. The stuff under {include,lib}/runtime is totally firewalled from the compiler and tiny (<6kB, though no attention has gone into optimizing that size). For example, we don't link in libSupport into the runtime, instead having our own bare bones replacements for basics like ArrayRef (the JITRuntime helps with bridging that gap, since it can depend on all common LLVM utilities). The overall features of npcomprt is that it exposes a module that with multiple function entry points. Each function has arguments and results that are tensor-valued, and npcomprt::Tensor is the runtime type that is used to interact with that (and a npcomprt::Ref<T> reference-counting wrapper is provided to wrap npcomprt::Tensor in the common case). From an implementation perspective, an npcomprt module at the LLVM/object/binary level exposes a single module descriptor struct that has pointers to other metadata (currently just a list of function metadata descriptors). All interactions with the npcomp runtime are keyed off of that module descriptor, including function lookups and dispatching. This is done to dodge platform ABI issues and also allow enough reflection to e.g. verify provided arguments. Most of the compiler-side work here was in LowerToNpcomprtABI and LowerToLLVM. Also, - Rename npcomp_rt/NpcompRt to npcomprt/Npcomprt; it was getting annoying to type the underscores/caps. - misc improvements to bash_helpers.sh	2020-07-08 19:36:19 -07:00
Stella Laurenzo	5aa2f0f9f6	Add a trivial copy elision canonicalization on ndarray->tensor. * This elides the very common code the compiler adds for chaining otherwise tensor-related numpy ops together. * More aggressive canonicalizations would require more advanced analysis.	2020-07-05 18:09:43 -07:00
Stella Laurenzo	fae15ec5e7	Allow the ndarray type to carry a shape.	2020-07-05 17:34:03 -07:00
Stella Laurenzo	046751254f	Refactor old tracing tests and remove deprecated ops. * Old doctests to run under lit. * Old custom filecheck tests -> pytest directory (under lit). * Rename some old ufunc ops in the tracer.	2020-06-29 16:19:03 -07:00
Stella Laurenzo	b2708e4687	Add test case for !numpy.ndarray.	2020-06-28 17:41:21 -07:00
Stella Laurenzo	7bd5733d38	Add "template function" ops and importer code. * This starts to lay down the infra for reasoning about calls * Adds the importer code to generate IR for function calls of compiler recognized static functions.	2020-06-26 18:36:36 -07:00
Stella Laurenzo	529873d13c	Wire up IREE compilation and runtime in a new backend test. * Adds python bindings for invoking flow, HAL, and VM lowering pipelines. * Adds pythong bindings for translating to VM module flatbuffer. * Adds a new backend_test/iree directory and configure lit to find the IREE python rt bindings. * Open code a simple_invoke.py that exercises the whole pipeline (need real APIs for a lot of this). * Fails when invoking the function because I never implemented argument marshaling for scalars :( * Plenty of stuff to do tomorrow.	2020-06-19 00:30:34 -07:00
Stella Laurenzo	308a54c3d0	Bump llvm-project to 52cae05e087b3d4fd02849fc37c387c720055ffb (2020/6/10). * Fixes compile errors from upstream. * XFAIL several tests that are now failing to legalize (will hand off to Sean).	2020-06-11 16:10:05 -07:00
Stella Laurenzo	8280b86c05	Aggregate all lit test targets under check-npcomp.	2020-06-07 14:35:58 -07:00
Stella Laurenzo	af4466197e	Add lit test suite for python compiler. * Adds a test for simple constants and fixes issues.	2020-06-07 14:29:39 -07:00
Sean Silva	92e45703ad	Remove XFAIL. This test seems to be passing, after a clean rebuild of everything (including MLIR).	2020-06-03 20:52:16 -07:00
Sean Silva	cd7258dbd4	Enable warnings by default. The secret here is LLVM_ENABLE_WARNINGS=ON. I also fixed a couple warnings, which gets us to be warning-clean. I noticed also that npcomp-run-mlir/basic.mlir seems to be failing. Maybe something since the latest integrate. My next commit (introduce npcomp mini runtime) will largely rewrite it though, so it'll get fixed then.	2020-06-03 20:39:34 -07:00
Sean Silva	e62e2e2915	Rename tests that run e2e-lowering-pipeline. This makes it more clear that they are grouped together. A directory seemed too heavyweight.	2020-06-01 19:34:44 -07:00
Sean Silva	7b9f0c3364	Add ability to run without optimizations. The default is to only do the bare minimum needed for correctness, since that stresses the layering of the system maximally.	2020-06-01 19:33:59 -07:00
Sean Silva	e8b1a07ef4	Initial NpcompRt (npcomp_rt) dialect boilerplate.	2020-06-01 19:07:53 -07:00
Sean Silva	e7b5a2b8a3	Make LowerRankedShapes clean up shape.from_extents ops. We were previously relying on a later canonicalization pass to clean them up, but it is a cleaner invariant if the pass gets rid of them itself.	2020-05-29 18:00:35 -07:00
Sean Silva	ccd5754b88	Rename `check-npcomp-opt` to just `check-npcomp`. It runs npcomp-run-mlir as well now, so having `-opt` in the name is confusing.	2020-05-29 16:12:10 -07:00
Sean Silva	ea822968fa	Add bare-bones npcomp-run-mlir. The code isn't super clean, but is a useful incremental step establishing most of the boilerplate for future enhancements. We can't print or return tensors yet so correctness TBD, but I've stepped into the running code in the debugger so I know it definitely is running. This is the first step to building out an npcomp mini-runtime. The mini-runtime doesn't have to be fancy or complex, but it should at least be layered nicely (which this code and the current compiler interaction with the "runtime" code is not). Now that we have boilerplate for e2e execution in some form, we can build that out.	2020-05-28 18:37:11 -07:00
Sean Silva	3a09455540	Use upstream shape.from_extents Replace our local `tcp.shape_from_extents` op with the upstream `shape.from_extents` op.	2020-05-21 14:51:01 -07:00
Sean Silva	1d3dbd9d5c	Lower to LLVM dialect. With this commit, we finish conversion to LLVM dialect, and should be ready for subsequent commits to convert to an LLVM module and let LLVM codegen to native machine code. This required a custom "lower to LLVM" pass to support lowering tcp.abort_if to a runtime call. In the future, this pass will grow to do type conversions for our own runtime types as we add those.	2020-05-20 18:56:10 -07:00
Sean Silva	be1971c4fc	Rename tcp.abort_if to tcp.shape_observe_error This more clearly captures its semantics as a structural "observer" of code that we currently mark as NoSideEffect but eventually lowers to eager error handling code. Also, update LowerRankedShapes to erase it, now that the layering here is clear. That pass reifies the eager error handling code, so the need for the dummy op to keep things alive isn't needed. With this change, we are now ready to start lowering to LLVM! This is the current print-ir-after-all from e2e-lowering-pipeline: https://reviews.llvm.org/P8221	2020-05-18 13:38:47 -07:00
Sean Silva	836a8d4bec	Lower tcp.alloc_memref ops to tcp.get_extent + std.alloc. - tcp.get_extent will be liminated while lowering shapes - std.alloc is supported by the upstream LLVM lowering.	2020-05-18 12:53:31 -07:00
Sean Silva	993338a12d	Lower to the upstream memref ABI. Specifically, we use unranked memrefs which get passed as a fixed-size set of arguments/returns. One big caveat about this is that returning results isn't going to work. See TODO in LowerTensorLoadOp. This is far from enough runtime-wise, but it starts to demarcate a plausible layering. Notice for example how this removes the runtime-dependence from LowerRankedShapes. Eventually, we want to have an `npcomp_rt` or `npcomp_hal` dialect with its own set of runtime types that will supercede this. See comments in LowerTensorLoadOp for more direction about where this is going to evolve.	2020-05-15 17:19:57 -07:00
Sean Silva	1b48d0d80b	Remove the present tcp.island. The idea was half-baked and after some deep thought felt like a solution looking for a problem. What we had here (and is removed in this patch) just wasn't pulling its weight. I cannot think of anything we would want to do with tcp.island as it is removed here beyond just sinking and merging them within a basic block, such that the witness argument is kind of pointless (only matters for hoisting). TCP compute ops like tcp.add and tcp.broadcast_to have the strong invariant of "pure or undefined behavior", which means they are always safe to sink. The island concept as removed here conferred no benefit. Also, I'll note that "islands" are a trick you can only play once in a system (unless they strictly nest). I have some early-stage thoughs on having an island concept that helps with modeling tensor shapes robustly which seems promising (the island would serve a similar role as tie_shape).	2020-05-14 15:19:37 -07:00
Sean Silva	889fe0d6c2	Tidy up test/E2E - Make rank1.mlir be the new "basic.mlir", as it is really the simplest case. - Move basic.mlir to mixed-ranks.mlir - Delete starting-from-linalg.mlir, it wasn't really useful anymore.	2020-05-14 14:59:55 -07:00
Sean Silva	eaeb4011e6	Lower !shape.shape to SSA values. This uses an approach inspired by what is done in IREE. See comments on LowerRankedShapes.cpp for how it works. The basic gist is that we have an op that creates a !shape.shape from a set of SSA values representing the extents, and then iteratively replace any op producing a !shape.shape with instances of that op.	2020-05-13 17:20:23 -07:00
Sean Silva	f525d4dbcf	Add custom assembly format for tcp.alloc_memref/tcp.get_extent This makes the IR a bit easier to scan.	2020-05-11 15:28:34 -07:00
Sean Silva	53c17dbed9	"Finish" tensor -> memref conversion. There's a lot of details to flesh out here, but the basic approach seems promising (see comments in createE2ELoweringPipeline). This approach will be put to the test when we try to do our first fusions since that tickles some of the nasty phase ordering issues involved here. But we're not there yet.	2020-05-11 15:00:12 -07:00
Sean Silva	e29aef855b	Initial TCF/TCP E2E seed. Very much WIP. This is enough to get tcf.add down to approximately the "linalg.generic on buffers" level of abstraction. (but there are nuances)	2020-05-08 20:20:41 -07:00
Stella Laurenzo	bc5ef81d68	Add basicpy.SlotObject type and ops to create/index into it. * This is intended to provide low-level modeling for built-in objects. * It is now possible to trace slice tuples (which are tuples of NoneType\|EllipsisType\|SlotObjectType<slice, ...>).	2020-05-05 18:16:01 -07:00
Stella Laurenzo	d3632af675	Add !numpy.any_dtype dialect type.	2020-04-29 18:20:42 -07:00
Stella Laurenzo	b4425fe1d2	Add numpy.ufunc_call op.	2020-04-29 17:49:56 -07:00
Stella Laurenzo	e845db8a20	Add builtin_ufunc and generic_ufunc ops.	2020-04-28 23:51:54 -07:00
Stella Laurenzo	953ef89a30	Add npcomp-opt and lit runner.	2020-04-26 17:55:15 -07:00

... 12 13 14 15 16

758 Commits (0a2d21b108602d2b11c208ca1a713a72f483f6c1)