SPIRV-Tools

mirror of https://github.com/KhronosGroup/SPIRV-Tools synced 2024-11-23 12:10:06 +00:00

Author	SHA1	Message	Date
Stephen McGroarty	e354984b09	Unroller support for multiple induction variables Support for multiple induction variables within a loop and support for loop condition operands <= and >=.	2018-02-27 11:50:08 +00:00
Steven Perron	94af58a350	Clean up variables before sroa In some shaders there are a lot of very large and deeply nested structures. This creates a lot of work for scalar replacement. Also, since commit `ca4457b` we have been very aggressive as rewriting variables. This has causes a large increase in compile time in creating and then deleting the instructions. To help low the costs, I want to run a cleanup of some of the easy loads and stores to remove. This reduces the number of symbols sroa has to work on. It also reduces the amount of code the simplifier has to simplify because it was not generated by sroa. To confirm the improvement, I ran numbers on three different sets of shaders: Time to run --legalize-hlsl: Set #1: 55.89s -> 12.0s Set #2: 1m44s -> 1m40.5s Set #3: 6.8s -> 5.7s Time to run -O Set #1: 18.8s -> 10.9s Set #2: 5m44s -> 4m17s Set #3: 7.8s -> 7.8s Contributes to #1328.	2018-02-22 21:40:58 -05:00
Steven Perron	3f19c2031a	Preserve analysies in the simplification pass Fixes a bug at the same time. In `UpdateDefUse`, if the definition already exists, we are not suppose to analyse it again. When you do the entries for the definition are deleted, and we don't want that. The check for this was wrong.	2018-02-22 16:06:30 -05:00
Lei Zhang	432dc40412	Appveyor: remove VS2015 configuration to reduce build time We already have VS2013 and VS2017, which should be good guards.	2018-02-22 15:17:04 -05:00
GregF	46a9ec9d23	Opt: Check for side-effects in DCEInst() This function now checks for side-effects before adding operand instructions to the dead instruction work list. Because this fix puts more pressure on IsCombinatorInstruction() to be correct, this commit adds all OpConstant* and OpType* instructions to combinator_ops_ set. Fixes #1341.	2018-02-22 12:24:13 -05:00
Alan Baker	01760d2f0f	Fixes #1338 . Handle OpConstantNull in branch/switch conditions * No longer assume the branch/switch condition must be bool or int constants (respectively) * Added a couple unit tests for each case	2018-02-21 10:22:39 -05:00
Steven Perron	51ecc7318f	Reduce instruction create and deletion during inlining. When inlining a function call the instructions in the same basic block as the call get cloned. The clone is added to the set of new blocks containing the inlined code, and the original instructions are deleted. This PR will change this so that we simply move the instructions to the new blocks. This saves on the creation and deletion of the instructions. Contributes to #1328.	2018-02-21 09:50:47 -05:00
Steven Perron	c1b936637e	Add Insert-extract elimination back into legalization passes. Fixes #1326.	2018-02-21 09:46:51 -05:00
Arseny Kapoulkine	309be423cc	Add folding for redundant add/sub/mul/div/mix operations This change implements instruction folding for arithmetic operations that are redundant, specifically: x + 0 = 0 + x = x x - 0 = x 0 - x = -x x * 0 = 0 * x = 0 x * 1 = 1 * x = x 0 / x = 0 x / 1 = x mix(a, b, 0) = a mix(a, b, 1) = b Cache ExtInst import id in feature manager This allows us to avoid string lookups during optimization; for now we just cache GLSL std450 import id but I can imagine caching more sets as they become utilized by the optimizer. Add tests for add/sub/mul/div/mix folding The tests cover scalar float/double cases, and some vector cases. Since most of the code for floating point folding is shared, the tests for vector folding are not as exhaustive as scalar. To test sub->negate folding I had to implement a custom fixture.	2018-02-20 18:29:27 -05:00
Steven Perron	fa3ac3cc33	Revert "Preserve analysies in the simplification pass" This reverts commit `ec3bbf093e`.	2018-02-20 18:21:25 -05:00
Steven Perron	ec3bbf093e	Preserve analysies in the simplification pass Building the def-use chains is very expensive, so we do not want to invalidate them it if is not necessary. At the moment, it seems like most optimizatoins are good at not invalidating the def-use chains, but simplification does. This PR get the simlification pass to keep the analysies valid. Contributes to #1328.	2018-02-20 14:45:08 -05:00
Diego Novillo	6c75050136	Speed up Phi insertion. On some shader code we have in our testsuite, Phi insertion is showing massive compile time slowdowns, particularly during destruction. The specific shader I was looking at has about 600 variables to keep track of and around 3200 basic blocks. The algorithm is currently O(var x blocks), which means maps with around 2M entries. This was taking about 8 minutes of compile time. This patch changes the tracking of stored variables to be more sparse. Instead of having every basic block contain all the tracked variables in the map, they now have only the variables actually stored in that block. This speeds up deallocation, which brings down compile time to about 1m20s. Note that this is not the definite fix for this. I will re-write Phi insertion to use a standard SSA rewriting algorithm (https://github.com/KhronosGroup/SPIRV-Tools/issues/893). This contributes to https://github.com/KhronosGroup/SPIRV-Tools/issues/1328.	2018-02-20 12:04:06 -05:00
Lei Zhang	ac9ab0dacf	Travis: require MacOS to build and test again	2018-02-20 11:25:38 -05:00
Steven Perron	9d95a91a9f	Fix folding insert feeding extract I mixed up two cases when folding an OpCompositeExtract that is feed by and OpCompositeInsert. The specific cases are demonstracted in the new test. I mixed up the conditions for the cases, and treated one like the other. Fixes #1323.	2018-02-20 11:22:51 -05:00
Alan Baker	c3f34d8bf3	Fixes #1300 . Adding checks for bad CCP transitions and unsettled values * Now track propagation status and assert on bad statuses * Added helper methods to access instruction propagation status * Modified the phi meet operator to properly reflect the paper it is based on * Modified SSA edge addition so that all edge are added, but only on state changes * Fixed a bug in instruction simulation where interesting conditional branches would not mark the interesting edge as executed * Added a test to catch this bug * Added an ostream operator for SSAPropagator::PropStatus	2018-02-18 19:41:34 -05:00
Andrew Woloszyn	e543b195df	Removed warnings from hex_float.h Bitcasting FloatProxy<->uint_type was hitting a warning with g++8.0.1. Replace bitcasts with new casting traits for FloatProxy.	2018-02-16 21:15:51 -05:00
Steven Perron	04cd63e5b9	Make better use of simplification pass The simplification pass works better after all of the dead branches are removed. So swapping them around in the legalization passes. Also adding the simplification pass to performance passes right after dead branch elimination. Added CCP to the legalization passes so we can propagate the constants into the branchs, and remove as many branches a possible. CCP is designed to still get opportunities even if the branches are dead, so it is a good place for it. Fixes #1118	2018-02-16 20:46:49 -05:00
Arseny Kapoulkine	1054413600	Add constant folding rules for floating-point comparison This change handles all 6 regular comparison types in two variations, ordered (true if values are ordered and comparison is true) and unordered (true if values are unordered or comparison is true). Ordered comparison matches the default floating-point behavior on host but we use std::isnan to check ordering explicitly anyway. This change also slightly reworks the floating-point folding support code to make it possible to define a folding operation that returns boolean instead of floating point. These tests exhaustively test ordered/unordered comparisons for float/double. Since for NaN inputs the comparison result doesn't depend on the comparison function, we just test == and !=; NaN inputs result in true unordered comparisons and false ordered comparisons.	2018-02-16 20:41:22 -05:00
Arseny Kapoulkine	27d23a92a0	Remove constants from constant manager in KillInst Registering a constant in constant manager establishes a relation between instruction that defined it and constant object. On complex shaders this could result in the constant definition getting removed as part of one of the DCE pass, and a subsequent simplification pass trying to use the defining instruction for the constant. To fix this, we now remove associated constant entries from constant manager when killing constant instructions; the constant object is still registered and can be remapped to a new instruction later. GetDefiningInstruction shouldn't ever return nullptr after this change so add an assertion to check for that.	2018-02-16 20:37:12 -05:00
Steven Perron	50f307f889	Simplify OpPhi instructions referencing unreachable continues In dead branch elimination, we already recognize unreachable continue blocks, and update OpPhi instruction accordingly. This change adds an extra check: if the head block has exactly 1 other incoming edge, then replace the OpPhi with the value from that edge. Fixes #1314.	2018-02-16 18:58:03 -05:00
Steven Perron	3756b387f3	Get CCP to use the constant floating point rules. Fixes #1311	2018-02-16 13:49:47 -05:00
Lei Zhang	efe286cd32	SubgroupBallotKHR can enable SubgroupSize & SubgroupLocalInvocationId	2018-02-16 10:02:18 -05:00
Nerijus Baliūnas	5d442fad2f	fix typo	2018-02-15 12:00:02 -05:00
David Neto	d2c0fce361	Invoke cmake via CMAKE_COMMAND variable Need to do this in case cmake is not on the path. This should fix the Android NDK build, as in when building the NDK itself.	2018-02-15 11:34:50 -05:00
Lei Zhang	f3a10470d3	Avoid using static unordered_map (#1304 ) unordered_map is not POD. Using it as static may cause problems when operator new() and operator delete() is customized. Also changed some function signatures to use const char* instead of std::string, which will give caller the flexibility to avoid creating a std::string.	2018-02-15 10:19:15 -05:00
Arseny Kapoulkine	32a8e04c7d	Add folding of redundant OpSelect insns We can fold OpSelect into one of the operands in two cases: - condition is constant - both results are the same Even if the original shader doesn't have either of these, if-conversion pass sometimes ends up generating instructions like %7127 = OpSelect %int %3220 %7058 %7058 And this optimization cleans them up.	2018-02-15 10:03:22 -05:00
Steven Perron	0e9f2f948a	Add id to name map Adding a map from an id to it set of OpName and OpMemberName instructions. This will be used in KillNameAndDecorates to kill the names for the ids that are being removed. In my test, the compile time for 50 shaders went from 1m57s to 55s. This was on linux using the release build. Fixes #1290.	2018-02-14 15:53:13 -05:00
Steven Perron	6669d8163d	Fold binary floating point operators. Adds the floating rules for FAdd, FDiv, FMul, and FSub. Contributes to #1164.	2018-02-14 15:48:15 -05:00
Stephen McGroarty	dd8400e150	Initial support for loop unrolling. This patch adds initial support for loop unrolling in the form of a series of utility classes which perform the unrolling. The pass can be run with the command spirv-opt --loop-unroll. This will unroll loops within the module which have the unroll hint set. The unroller imposes a number of requirements on the loops it can unroll. These are documented in the comments for the LoopUtils::CanPerformUnroll method in loop_utils.h. Some of the restrictions will be lifted in future patches.	2018-02-14 15:44:38 -05:00
Alan Baker	229ebc0665	Fixes #1295 . Mark undef values as varying in ccp. * Undef now marked as varying in ccp * this prevents incorrect meet operations since phis were always not interesting * added a test to catch the bug	2018-02-14 10:21:26 -05:00
Diego Novillo	08699920ad	Cleanup. Use proper #include guard. NFC.	2018-02-12 13:21:48 -05:00
Steven Perron	06b437dedc	Avoid using the def-use manager during inlining. There seems to only be a single location where the def-use manager is used. It is to get information about a type. We can do that with the type manager instead. Fixes #1285	2018-02-12 09:47:55 -05:00
Arseny Kapoulkine	70bf3514e8	Fix spirv.h include to rely on include paths This is important when SPIRV-Headers are not checked out to external/ folder and mirrors other places in the code where spirv.h is included.	2018-02-09 18:29:17 -08:00
Steven Perron	1d7b1423f9	Add folding of OpCompositeExtract and OpConstantComposite constant instructions. Create files for constant folding rules. Add the rules for OpConstantComposite and OpCompositeExtract.	2018-02-09 17:52:33 -05:00
David Neto	886859159e	Fix generation of Vim syntax file	2018-02-09 17:47:51 -05:00
Józef Kucia	4e4a254bc8	Do not hardcode libdir and includedir in pkg config files	2018-02-09 12:43:03 +01:00
Steven Perron	1a849ffb60	Add header files missing from CMakeLists.txt	2018-02-08 23:02:22 -05:00
Alexander Johnston	84ccd0b9ae	Loop invariant code motion initial implementation	2018-02-08 22:55:47 -05:00
GregF	ca4457b4b6	SROA: Do replacement on structs with no partial references.	2018-02-08 15:20:02 -05:00
Steven Perron	06cdb96984	Make use of the instruction folder. Implementation of the simplification pass. - Create pass that calls the instruction folder on each instruction and propagate instructions that fold to a copy. This will do copy propagation as well. - Did not use the propagator engine because I want to modify the instruction as we go along. - Change folding to not allocate new instructions, but make changes in place. This change had a big impact on compile time. - Add simplification pass to the legalization passes in place of insert-extract elimination. - Added test cases for new folding rules. - Added tests for the simplification pass - Added a method to the CFG to apply a function to the basic blocks in reverse post order. Contributes to #1164.	2018-02-07 23:01:47 -05:00
Andrey Tuganov	a61e4c1356	Disable check which fails Vulkan CTS	2018-02-07 13:31:35 -05:00
Andrey Tuganov	2f0c3aaa11	Add Vulkan-specific validation rules for atomics Added atomic instructions validation rules from https://www.khronos.org/registry/vulkan/specs/1.0/html/vkspec.html#spirvenv-module-validation	2018-02-07 13:31:35 -05:00
Józef Kucia	3013897556	Build SPIRV-Tools as shared library Add pkg-config file for shared libraries Properly build SPIRV-Tools DLL Test C interface with shared library Set PATH to shared library file for c_interface_shared test Otherwise, the test won't find SPIRV-Tools-shared.dll. Do not use private functions when testing with shared library Make all symbols hidden by default for shared library target	2018-02-07 10:43:32 -05:00
David Neto	b1c9c4e8c0	Enable Visual Studio 2013 again Disable use of Effcee and RE2 with MSVC compilers older than Visual Studio 2015 since RE2 doesn't support them.	2018-02-06 14:40:28 -05:00
David Neto	e7fafdaa68	Fix test inclusion when Effcee is absent	2018-02-06 12:10:50 -05:00
David Neto	c452bfd054	Update CHANGES	2018-02-06 11:47:44 -05:00
Alan Baker	871022772e	Registering a type now rebuilds it out of memory owned by the manager. * Added TypeManager::RebuildType * rebuilds the type and its constituent types in terms of memory owned by the manager. * Used by TypeManager::RegisterType to properly allocate memory * Adding an unit test to expose the issue * Added some tests to provide coverage of RebuildType * Added an accessor to the target pointer for a forward pointer	2018-02-06 10:17:56 -05:00
GregF	860b2ee5fc	ADCE: Fix combinator initialization The combinator initialization was only looking at the capabilities in the shader and not the inferred capabilities. Geometry and tessellation shaders were not setting the Shader capability which is inferred. So the combinator set was not initialized correctly causing problems for ADCE.	2018-02-05 16:54:03 -05:00
David Neto	9e19fc0f31	VS2013: LoopDescriptor LoopContainerType can't contain unique_ptr The loop descriptor must explicitly manage the storage for contained Loop objects. Fixes #1262	2018-02-05 14:19:21 -05:00
Andrey Tuganov	12e6860d07	Add barrier instructions validation pass	2018-02-05 13:14:55 -05:00

1 2 3 4 5 ...

1312 Commits